2606.27814v6 Jun 26, 2026 cs.AI

ATOD: 어닐링된 턴 인식 온-폴리시 증류를 이용한 다중 턴 에이전트 작업

ATOD: Annealed Turn-Aware On-Policy Distillation for Multi-Turn Agentic Tasks

Zefang Zong
Zefang Zong
Citations: 747
h-index: 12
Mo Li
Mo Li
Tsinghua University
Citations: 964
h-index: 6
Qitai Tan
Qitai Tan
Citations: 8
h-index: 2
Peng Chen
Peng Chen
Citations: 61
h-index: 3
Yipeng Shi
Yipeng Shi
Citations: 0
h-index: 0
Yang Li
Yang Li
Citations: 8
h-index: 2

장기적인 상호 작용 작업을 위한 소형 언어 모델 에이전트를 학습하려면 빠른 모방과 보상 기반 개선이 모두 필요합니다. 온-폴리시 증류(OPD)는 풍부한 교사 가이드를 제공하며 일반적으로 초기 단계에서 빠르게 성능이 향상되지만, 학생 모델이 교사 모델에 가까워지면 성능 향상이 포화되어 최종 성능의 한계가 발생합니다. 강화 학습(RL)은 환경 보상을 직접 최적화하여 더 높은 보상 수준으로 탐색적인 개선을 장려하지만, 희소하고 지연된 피드백은 초기 단계 학습 효율성을 OPD보다 훨씬 떨어지게 만듭니다. 본 논문에서는 이러한 상호 보완성을 명시적으로 활용하는 하이브리드 온라인 증류 알고리즘인 ATOD(Annealed Turn-aware On-policy Distillation)를 제안합니다. (1) ATOD는 어닐링된 OPD-RL 스케줄을 사용합니다. OPD는 초기 학습 단계에서 교사 수준의 동작에 도달하는 데 주도적인 역할을 하며, RL은 점진적으로 강화되어 보상 기반 탐색을 촉진합니다. (2) ATOD는 턴 수준의 불일치-불확실성 재가중(T-DUR)을 도입하여 긴 시퀀스에서 높은 불일치 또는 불확실성을 보이는 턴에 우선순위를 부여하기 위해 증류 신호를 소프트하게 제어합니다. ALFWorld, WebShop 및 Search-QA 데이터셋에서의 실험 결과, ATOD는 경쟁적인 사후 학습 기준 모델보다 일관되게 우수한 성능을 보입니다. 세 가지 학생 모델 크기에서 ATOD는 평균 성공률을 OPD보다 4.16 포인트, GRPO보다 23.62 포인트 향상시켰으며, 해당 교사 모델보다 2.16 포인트 더 높은 성능을 달성했습니다.

Original Abstract

Training small language-model agents for long-horizon interactive tasks requires both fast imitation and reward-driven improvement. On-policy distillation (OPD) provides dense teacher guidance and typically improves rapidly in the early stage, but its gains saturate once the student approaches the teacher, limiting the final performance ceiling. Reinforcement learning (RL) directly optimizes environment rewards and encourages exploratory improvement toward a higher reward-defined ceiling, but sparse and delayed feedback makes early-stage learning much less efficient than OPD. In this paper, we propose ATOD (Annealed Turn-aware On-policy Distillation), a hybrid online distillation algorithm that explicitly exploits this complementarity. (1) ATOD uses an annealed OPD-RL schedule: OPD dominates early training to approach teacher-level behavior, while RL is gradually strengthened to drive reward-based exploration. (2) ATOD introduces Turn-level Disagreement-Uncertainty Reweighting (T-DUR), which softly gates the distillation sig- nal to prioritize turns with high disagreement or uncertainty in long trajectories. Experiments on ALFWorld, WebShop, and Search-QA show that ATOD consistently outperforms competing post-training baselines: across the three student sizes, ATOD improves average success rate by 4.16 points over OPD and 23.62 points over GRPO, while surpassing the corresponding teacher models by 2.16 points.

1 Citations
0 Influential
6 Altmetric
31.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!