2608.05219v1 Aug 05, 2026 cs.AI

특권적인 지도가 일치하지 않을 때: 상태 매칭 라우팅 및 맥락화된 자기 증류를 통한 다중 회선 에이전트

When Privileged Guidance Misaligns: State-Matched Routing and Contextualized Self-Distillation for Multi-Turn Agents

Jun Ling
Jun Ling
Citations: 3
h-index: 1
Junzhuo Liu
Junzhuo Liu
University of Electronic Science and Technology of China
Citations: 35
h-index: 3
Peng Wang
Peng Wang
Citations: 40
h-index: 3
Weiwei Li
Weiwei Li
University of Electronic Science and Technology of China
Citations: 26
h-index: 2

프라이빌레지드 온폴리시 증류는 동기화된 교사가 학습 전용 참조 자료(예: 성공적인 경로)에 접근하여 모든 단계에서 학생의 응답을 재평가함으로써 다중 회선 에이전트에 대한 밀집적 감독 신호를 제공합니다. 그러나 상호 작용 환경에서는 학생의 이전 행동이 지속적으로 실행 상태를 변경합니다. 학생이 다른 행동을 취하거나 하위 목표를 다른 순서로 완료하면, 롤아웃 과정에서 참조에 포함되지 않은 상태에 도달할 수 있으며, 이는 실제로 도달한 상태에 대한 신뢰할 수 없는 지침원이 됩니다. 따라서 프라이빌레지드 증류를 무분별하게 적용하는 것은 상태-참조 불일치를 야기합니다. 이러한 불일치는 핵심 목표를 제시합니다: 학생의 현재 실행 상태와 호환되는 프라이빌레지드 참조 지침을 제공하는 것입니다. 우리는 상태 매칭 라우팅 및 맥락화된 자기 증류 (SMRC-SD)를 도입하며, 이는 프라이빌레지드 경로가 온폴리시 학생을 언제 그리고 어떻게 안내해야 하는지를 명시적으로 결정합니다. 각 단계에서 SMRC-SD는 학생의 현재 실행 상태가 참조 경로에 있는 지원되는 상태와 일치하는지 확인합니다. 증류는 일치하는 상태에서만 적용되며, 참조가 지역적으로 호환 가능한 지침을 제공하지 않는 단계를 필터링합니다. 또한, SMRC-SD는 각 일치하는 상태에 대해 성공적인 경로로부터 상태 의존적인 교사 컨텍스트를 구성하여, 실제로 도달한 상태에 기반한 감독 신호를 제공합니다. ALFWorld 및 WebShop에서 SMRC-SD는 조건 없는 성공적인 전체 경로 증류보다 꾸준히 우수한 성능을 보입니다. Qwen3-1.7B 모델로 사용할 때, ALFWorld에서 작업 성공률이 $0.746$에서 $0.865$로, WebShop에서 $0.574$에서 $0.693$으로 향상되었습니다. 제어된 라우팅 및 컨텍스트 제거 실험을 통해 지역적으로 지원되는 단계를 선택하고 상태와 호환 가능한 교사 컨텍스트를 구성하는 것이 이러한 성능 향상의 기여 요인임을 확인했습니다. 코드는 https://github.com/liujunzhuo/SMRC-SD 에서 확인할 수 있습니다.

Original Abstract

Privileged on-policy distillation provides dense supervision for multi-turn agents by allowing a synchronized teacher to re-score the student's response at every turn with access to training-only references, such as successful trajectories. In interactive environments, however, the student's preceding actions continually change the execution state. As the student takes different actions or completes subgoals in a different order, its rollout may reach states not covered by the reference, making the reference an unreliable source of guidance for the state actually reached. Applying privileged distillation indiscriminately therefore creates state--reference mismatch. This mismatch motivates a central objective: providing privileged reference guidance that remains compatible with the student's current execution state. We introduce State-Matched Routing and Contextualized Self-Distillation (SMRC-SD), which explicitly determines when and how a privileged trajectory should guide an on-policy student. At each turn, SMRC-SD verifies whether the student's current execution state matches a supported state along the reference trajectory. Distillation is applied only at matched states, filtering out turns for which the reference lacks locally compatible guidance. For each matched state, SMRC-SD further constructs state-conditioned teacher context from the successful trajectory, grounding supervision in the state actually reached. Across ALFWorld and WebShop, SMRC-SD consistently outperforms unconditional successful full-path distillation. With Qwen3-1.7B, it improves task success from $0.746$ to $0.865$ on ALFWorld and from $0.574$ to $0.693$ on WebShop. Controlled routing and context ablations support both selecting locally supported turns and constructing state-compatible teacher context as contributors to these gains. Code is available at https://github.com/liujunzhuo/SMRC-SD.

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!