2605.28396v1 May 27, 2026 cs.LG

ADWIN: 수평 정보를 고려한 온라인 정책 증류를 위한 적응형 윈도우

ADWIN: Adaptive Windows for Horizon-Aware On-Policy Distillation

Clive Bai
Clive Bai
Citations: 10
h-index: 2
Saiyong Yang
Saiyong Yang
Citations: 104
h-index: 5
Weijie Liu
Weijie Liu
Citations: 94
h-index: 4
Kun Liang
Kun Liang
Citations: 13
h-index: 3
Chenming Tang
Chenming Tang
Citations: 21
h-index: 3
Yunfang Wu
Yunfang Wu
Citations: 26
h-index: 3

온라인 정책 증류(OPD)는 학생 모델이 생성한 경로에 대한 교사의 피드백을 활용하여 학습함으로써 추론 행동을 전송합니다. 그러나 표준적인 전체 시퀀스 학습 방식은 매 업데이트마다 비용이 많이 드는 완료 과정을 필요로 하며, 현재 학생 모델에게 낮은 한계 가치를 갖는 후반 단계에 과도하게 감독 신호를 할당할 수 있습니다. 본 논문에서는 유용한 감독 범위 관점에서 이 가정을 재검토합니다. 학생 모델이 생성한 시퀀스는 교사가 선호하는 경로에서 벗어날 수 있으며, 정렬된 접두사(prefix)는 이미 장기적인 OPD 업데이트 방향을 유지할 수 있습니다. 우리는 OPD를 위한 적응형 윈도우 프레임워크인 ADWIN을 제안합니다. ADWIN은 시퀀스 길이를 온라인 허용 여부 결정으로 취급하며, 짧은 교사 기반 접두사를 사용하여 학습하고, 지연된 전체 시퀀스 검증(probe)을 통해 접두사와 전체 시퀀스의 일관성을 확인하고, 노후화 제어를 통해 다음 범위를 조정합니다. 단일 작업, 다중 작업 및 강-약(strong-to-weak) 설정에서 수학 및 코드 추론 벤치마크를 사용하여 ADWIN은 전체 시퀀스 OPD 및 접두사 기반 모델보다 정확도와 계산 효율성 간의 균형을 개선하며, 최대 4.1배까지 엔드투엔드 학습 비용을 줄이는 동시에 유사하거나 더 높은 정확도를 달성합니다.

Original Abstract

On-policy distillation (OPD) transfers reasoning behavior by training a student on teacher feedback along student-generated trajectories, but standard full-rollout training ties every update to a costly completion and can over-allocate supervision to late positions with low marginal value for the current student. We revisit this assumption through the useful supervision horizon: student-induced rollouts can drift from teacher-preferred continuations, while aligned prefixes may already preserve the long-horizon OPD update direction. We propose ADWIN, an adaptive-window framework for OPD that treats rollout length as an online admissibility decision, training on short teacher-anchored prefixes while using delayed full-rollout probes to audit prefix--full alignment and adapt the next horizon with staleness control. Across math and code reasoning benchmarks in single-task, multi-task, and strong-to-weak settings, ADWIN improves the accuracy--compute trade-off over full-rollout OPD and prefix-based baselines, reducing end-to-end training cost by up to 4.1 times while achieving comparable or better accuracy.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!