2608.01953v1 Aug 03, 2026 cs.CL

증류하기 전에 미래를 예측하세요: 에이전트 기반 온라인 증류의 가이드라인 효과 검증을 위한 미래 경로 예측

Look Ahead Before You Distill: Future Trajectory Validation of Teacher Guidance for Agentic On-Policy Distillation

Xuyang Liu
Xuyang Liu
Citations: 260
h-index: 9
Linfeng Zhang
Linfeng Zhang
Citations: 76
h-index: 5
Lu Pan
Lu Pan
Citations: 38
h-index: 3
Junxi Wang
Junxi Wang
Citations: 22
h-index: 3
Delin Mao
Delin Mao
Citations: 0
h-index: 0
Hongbo Qiao
Hongbo Qiao
Citations: 0
h-index: 0
Chishui Chen
Chishui Chen
Citations: 1
h-index: 1
Yao Fan
Yao Fan
Citations: 0
h-index: 0
Chenghao Sun
Chenghao Sun
Citations: 31
h-index: 2
Chenxing Sun
Chenxing Sun
Citations: 0
h-index: 0
Zuowei Zhang
Zuowei Zhang
Citations: 0
h-index: 0
Te Sun
Te Sun
Citations: 129
h-index: 4
Yangen Hu
Yangen Hu
Citations: 23
h-index: 2
Y. Yang
Y. Yang
Citations: 32
h-index: 1

온라인 증류(OPD)는 학생 에이전트가 방문하는 상태에 대해 교사의 지도를 제공하여 학습과 추론 간의 분포 격차를 줄입니다. 그러나 다단계 에이전트 작업에서는 학생의 편차가 시간이 지남에 따라 누적되어 경로가 교사 지도가 효과적인 상태에서 멀어질 수 있습니다. 우리의 정량적 분석 결과는 불확실성이 높은 상태가 교사 지도를 위한 유망한 기회를 제공하지만, 그러한 지도가 실제로 이점 있는지 여부를 판단하려면 학생의 후속 경로에 미치는 영향을 검토해야 함을 보여줍니다. 우리는 FutureBridge-OPD (FTB)를 제안합니다. FTB는 불확실성이 높은 상태에서 짧은 교사 연결(teacher bridge)을 실행하고, 이를 통해 얻어진 학생의 이어지는 행동을 사용하여 해당 연결이 교사의 긍정적인 증류 신호의 밀도를 증가시키는지 여부를 평가합니다. ALFWorld, WebShop 및 ScienceWorld 데이터셋에서 Qwen3-32B 모델을 교사로, Qwen3-1.7B 모델을 학생으로 사용했을 때, FTB는 기존의 OPD와 TCOD 방법보다 각각 평균 16.6점과 7.6점 높았습니다. 또한 FTB는 다양한 규모의 학생 모델 및 교사 모델 환경에서도 효과적인 성능을 보였습니다. 저희 코드는 다음 주소에서 공개적으로 이용할 수 있습니다: https://github.com/ChenChiShui/FutureBridge-OPD.

Original Abstract

On-policy distillation (OPD) provides teacher supervision on states visited by the student, reducing the distribution gap between training and inference. However, in multi-turn agentic tasks, student deviations may accumulate over time, gradually moving the trajectory away from states where teacher guidance remains effective. Our quantitative analysis further shows that high-disagreement states offer promising opportunities for teacher guidance, but determining whether such guidance is beneficial requires examining its effect on subsequent student trajectories. We propose FutureBridge-OPD (FTB), which executes a short teacher bridge at a high disagreement state and uses the resulting student continuation to assess whether the bridge increases the density of positive distillation signals relative to the teacher. On ALFWorld, WebShop, and ScienceWorld, under the main Qwen3-32B teacher to Qwen3-1.7B student setting, FTB outperforms vanilla OPD and TCOD by an average of 16.6 and 7.6 points, respectively, and remains effective across student scales and teacher settings. Our code is publicly available at https://github.com/ChenChiShui/FutureBridge-OPD.

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!