2608.02110v1 Aug 03, 2026 cs.CL

IACM-RL: 의도 기반 컨텍스트 관리 및 강화 학습 - 복잡한 도구 실행 시 동적 의도 변화에 대한 접근

IACM-RL: Intent-Aware Context Management and Reinforcement Learning for Complex Tool Invocation under Dynamic Intent Fluctuations

Xuanjing Huang
Xuanjing Huang
Citations: 3,773
h-index: 33
Xipeng Qiu
Xipeng Qiu
Citations: 1,073
h-index: 15
Zhiheng Xi
Zhiheng Xi
Citations: 1,507
h-index: 17
Shihan Dou
Shihan Dou
Citations: 4,745
h-index: 26
Chenhao Huang
Chenhao Huang
Citations: 66
h-index: 4
Junjie Ye
Junjie Ye
Fudan University
Citations: 1,596
h-index: 17
Tao Gui
Tao Gui
Citations: 450
h-index: 7
Yunke Zhang
Yunke Zhang
Citations: 41
h-index: 3
Ming Zhang
Ming Zhang
Citations: 2,499
h-index: 11
Shichun Liu
Shichun Liu
Citations: 943
h-index: 10
Jiahang Lin
Jiahang Lin
Citations: 199
h-index: 3
Honglin Guo
Honglin Guo
Fudan University
Citations: 554
h-index: 10
Yunbin Zhao
Yunbin Zhao
Citations: 0
h-index: 0
Dingwei Zhu
Dingwei Zhu
Citations: 11
h-index: 2
Jiazheng Zhang
Jiazheng Zhang
Citations: 121
h-index: 6
Qi Zhang
Qi Zhang
Citations: 101
h-index: 4
Xin Guo
Xin Guo
Citations: 2,352
h-index: 6
Yuhui Wang
Yuhui Wang
Citations: 30
h-index: 3
Zhonghang Lu
Zhonghang Lu
Citations: 0
h-index: 0
Jiahan Li
Jiahan Li
Citations: 249
h-index: 2
Chengjun Pan
Chengjun Pan
Citations: 0
h-index: 0
Yunxian Yang
Yunxian Yang
Citations: 0
h-index: 0
Zhuohui Sheng
Zhuohui Sheng
Citations: 0
h-index: 0
Yajie Yang
Yajie Yang
Citations: 42
h-index: 3
Junlin Shang
Junlin Shang
Citations: 4
h-index: 1

실제 환경에서 장기적인 도구 실행은 사용자의 의도 변동으로 인해 심각한 어려움을 겪습니다. 기존 방법들은 암묵적인 히스토리 스캔이나 텍스트 압축을 통해 안정성을 확보하려고 시도하지만, 대부분 단순화된 시나리오를 가정하며 완벽한 지시 사항이 주어질 것이라고 전제합니다. 불가피하게, 변화하는 환경 속에서 오래된 제약 조건은 모델의 집중도를 희석시켜 파괴적인 의도 편차와 무한 API 루프를 유발할 수 있습니다. 이러한 문제를 해결하기 위해, 우리는 강력한 도구 실행을 위한 포괄적인 프레임워크인 IACM-RL을 제안합니다. 첫째, 우리는 13가지 세분화된 변동 시나리오를 포함하는 DynamicIntent 파이프라인을 구축하고, 이를 통해 얻어진 데이터를 기반으로 5차원 진단 지표 모음을 제공합니다. 둘째, IACM-RL은 BeliefState 기반의 자체 생성 컨텍스트 관리자를 사용하여 변화하는 목표를 능동적으로 추적하고, 구조적인 stale flags를 활용하여 덮어쓰여진 매개변수를 격리합니다. 이 상태 추적 기능을 자율적으로 내재화하기 위해, 우리는 계층적인 의도 기반 보상과 함께 세 가지 추가적인 손실 함수(행동 교정, CM 추출 및 상태 증류)를 사용하여 정책을 최적화했습니다. DynamicIntent, BFCL-V3 및 $ au^2$-Bench에 대한 실험 결과는 IACM-RL이 기존 방법보다 훨씬 우수한 성능을 보이며, 무한 루프와 오래된 컨텍스트 오류를 줄이고 일반화 능력을 향상시킨다는 것을 보여줍니다.

Original Abstract

Executing long-horizon tool invocations in real-world environments is severely challenged by dynamic user intent noise. Existing methods attempt robustness via implicit history scanning or text compression, yet predominantly assume perfect instructions in simplistic scenarios. Inevitably, under fluctuating contexts, obsolete constraints dilute model attention, triggering catastrophic intent deviation and infinite API loops. To resolve this, we propose IACM-RL, a comprehensive framework for robust tool invocation. First, we introduce the DynamicIntent pipeline, synthesizing trajectories across 13 fine-grained fluctuation scenarios, paired with a five-dimensional diagnostic metric suite. Second, IACM-RL deploys a BeliefState-based Self-Generated Context Manager that proactively tracks shifting goals and isolates overwritten parameters using structural stale flags. To autonomously internalize this state-tracking capability, we optimize the policy using a hierarchical intent-driven reward alongside three auxiliary losses (action calibration, CM extraction, and state distillation). Experiments on DynamicIntent, BFCL-V3, and $\mathrmτ^2$-Bench demonstrate that IACM-RL significantly outperforms baselines, reducing infinite loops and stale context errors while enhancing out-of-domain generalization.

0 Citations
0 Influential
16.5 Altmetric
82.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!