IACM-RL: 의도 기반 컨텍스트 관리 및 강화 학습 - 복잡한 도구 실행 시 동적 의도 변화에 대한 접근
IACM-RL: Intent-Aware Context Management and Reinforcement Learning for Complex Tool Invocation under Dynamic Intent Fluctuations
실제 환경에서 장기적인 도구 실행은 사용자의 의도 변동으로 인해 심각한 어려움을 겪습니다. 기존 방법들은 암묵적인 히스토리 스캔이나 텍스트 압축을 통해 안정성을 확보하려고 시도하지만, 대부분 단순화된 시나리오를 가정하며 완벽한 지시 사항이 주어질 것이라고 전제합니다. 불가피하게, 변화하는 환경 속에서 오래된 제약 조건은 모델의 집중도를 희석시켜 파괴적인 의도 편차와 무한 API 루프를 유발할 수 있습니다. 이러한 문제를 해결하기 위해, 우리는 강력한 도구 실행을 위한 포괄적인 프레임워크인 IACM-RL을 제안합니다. 첫째, 우리는 13가지 세분화된 변동 시나리오를 포함하는 DynamicIntent 파이프라인을 구축하고, 이를 통해 얻어진 데이터를 기반으로 5차원 진단 지표 모음을 제공합니다. 둘째, IACM-RL은 BeliefState 기반의 자체 생성 컨텍스트 관리자를 사용하여 변화하는 목표를 능동적으로 추적하고, 구조적인 stale flags를 활용하여 덮어쓰여진 매개변수를 격리합니다. 이 상태 추적 기능을 자율적으로 내재화하기 위해, 우리는 계층적인 의도 기반 보상과 함께 세 가지 추가적인 손실 함수(행동 교정, CM 추출 및 상태 증류)를 사용하여 정책을 최적화했습니다. DynamicIntent, BFCL-V3 및 $ au^2$-Bench에 대한 실험 결과는 IACM-RL이 기존 방법보다 훨씬 우수한 성능을 보이며, 무한 루프와 오래된 컨텍스트 오류를 줄이고 일반화 능력을 향상시킨다는 것을 보여줍니다.
Executing long-horizon tool invocations in real-world environments is severely challenged by dynamic user intent noise. Existing methods attempt robustness via implicit history scanning or text compression, yet predominantly assume perfect instructions in simplistic scenarios. Inevitably, under fluctuating contexts, obsolete constraints dilute model attention, triggering catastrophic intent deviation and infinite API loops. To resolve this, we propose IACM-RL, a comprehensive framework for robust tool invocation. First, we introduce the DynamicIntent pipeline, synthesizing trajectories across 13 fine-grained fluctuation scenarios, paired with a five-dimensional diagnostic metric suite. Second, IACM-RL deploys a BeliefState-based Self-Generated Context Manager that proactively tracks shifting goals and isolates overwritten parameters using structural stale flags. To autonomously internalize this state-tracking capability, we optimize the policy using a hierarchical intent-driven reward alongside three auxiliary losses (action calibration, CM extraction, and state distillation). Experiments on DynamicIntent, BFCL-V3, and $\mathrmτ^2$-Bench demonstrate that IACM-RL significantly outperforms baselines, reducing infinite loops and stale context errors while enhancing out-of-domain generalization.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.