2607.26057v1 Jul 28, 2026 cs.CL

지휘봉을 넘겨라: 경로 기반 온라인 정책 증류

Pass the Baton: Trajectory-Relayed On-Policy Distillation

Yongliang Shen
Yongliang Shen
Citations: 397
h-index: 10
Haiwen Hong
Haiwen Hong
Citations: 238
h-index: 5
Haolei Xu
Haolei Xu
Citations: 102
h-index: 4
Hongxing Li
Hongxing Li
Citations: 145
h-index: 4
Weiming Lu
Weiming Lu
Citations: 4,584
h-index: 30
Zixuan Ni
Zixuan Ni
Citations: 0
h-index: 0
Yiwen Qiu
Yiwen Qiu
Citations: 30
h-index: 4
Xiaowen Xu
Xiaowen Xu
Citations: 41
h-index: 1

온라인 정책 증류(OPD)는 학생 모델의 자체적인 추론 과정을 활용하여 토큰 수준의 지침을 제공하지만, '접두사 실패'라는 문제점을 가지고 있습니다. 즉, 학생 모델이 잘못된 추론 방향을 선택하면 이후 모든 생성 과정이 이 오류에 기반하게 되어 잘못된 결과를 초래하고 신뢰할 수 없는 지침을 제공하며 컴퓨팅 자원을 낭비합니다. 우리는 실패한 접두사에서 교사와 학생 모델 간의 추론 방향 불일치를 발견했습니다. 교사는 방향 전환을 시도하는 반면, 학생은 원래 방향을 유지하려는 경향이 있습니다. 이러한 불일치를 활용하여 '릴레이 온-라인 정책 증류(Relay-OPD)'라는 새로운 방법을 제안합니다. 훈련 과정에서 Relay-OPD는 특정 지점에서 교사 모델이 잠시 개입하여 추론 과정을 이어가는 '교사 구간'을 생성하고, 이후 학생 모델이 이 추론 과정을 이어받아 최적화하는 '릴레이 경로'를 구축합니다. 제한된 릴레이 예산은 중요한 초기 단계에서만 개입을 집중시키고 학생 모델의 정책으로부터 크게 벗어나지 않도록 합니다. Qwen3-4B-Instruct-2507 교사 모델과 Qwen3-0.6B/1.7B 학생 모델을 사용하여 8개의 수학적 추론 벤치마크에서 실험한 결과, Relay-OPD는 모든 벤치마크에서 최고 또는 두 번째로 높은 성능을 달성했습니다. 특히 1.7B 모델의 경우 평균적으로 표준 OPD보다 +5.73% 더 높고 FastOPD와 같은 강력한 기준 모델보다 +1.49% 더 높은 성능을 보였습니다. 또한 0.6B 모델에서도 일관된 성능 향상을 확인했으며, 훈련에 필요한 경로 길이를 50% 이상 단축했습니다.

Original Abstract

On-policy distillation (OPD) grounds token-level supervision in the student's own trajectory, yet suffers from prefix failure: once the student commits to a wrong reasoning direction, all subsequent generation builds on this deviation, producing misdirected continuations that elicit unreliable supervision and waste compute. We identify a teacher-student continuation asymmetry on failed prefixes, where the teacher tends to redirect while the student continues along the original direction, and convert it into a label-free handoff trigger in Relay On-Policy Distillation (Relay-OPD). During training, Relay-OPD constructs relay trajectories by letting the teacher briefly take over at detected trigger points to produce a teacher leg, after which the student resumes and is optimized on the resulting trajectory. A limited relay budget concentrates intervention on critical early positions while limiting departure from the student policy. With a Qwen3-4B-Instruct-2507 teacher and Qwen3-0.6B/1.7B-Non-Thinking students on eight mathematical reasoning benchmarks, Relay-OPD achieves the best or second-best results on every benchmark, outperforming standard OPD by +5.73% and the strongest baseline FastOPD by +1.49% on average for 1.7B, with consistent gains at 0.6B. Training trajectory length is reduced by over 50%.

1 Citations
0 Influential
15 Altmetric
76.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!