2607.29494v1 Jul 31, 2026 cs.LG

적응형 FastOPD: 효율적인 온-폴리시 증류를 위한 진행 상황 기반 롤아웃 호라이즌 확장

Adaptive FastOPD: Progress-Aware Rollout Horizon Expansion for Efficient On-Policy Distillation

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Qian Tan
Qian Tan
Citations: 214
h-index: 3
Lei Jiang
Lei Jiang
Citations: 116
h-index: 5
Huaifei Liang
Huaifei Liang
Citations: 0
h-index: 0

온-폴리시 증류(OPD)는 학생이 생성한 경로에 대한 풍부한 교사 모델의 지침을 제공하지만, 온라인 롤아웃 과정은 상당한 계산 비용을 발생시키는데, 특히 몇몇 긴 응답으로 인해 배치 완료가 지연되는 경우 더욱 그렇습니다. 기존의 가속화 방법들은 일반적으로 고정된 예산 또는 절대적인 교사-학생 동의 임계값을 사용하여 롤아웃 길이를 제어하는데, 이는 서로 다른 모델 및 학습 단계에서의 학습 진행 상황을 제대로 반영하지 못할 수 있습니다. 본 논문에서는 Adaptive FastOPD라는 진행 상황 기반 전략을 제안합니다. 이 방법은 학습이 현재 경계 영역 근처에서 정체되었고 현재 호라이즌이 충분히 활용되고 있을 때에만 롤아웃 범위를 확장합니다. 전자는 각 호라이즌에 진입할 때의 값과 비교하여 측정된 네 가지 교사-학생 신호로부터 결정되며, 이는 미리 정의된 단계 간격 또는 원시 동의 신호에 대한 절대 임계값보다 특정 단계별 진행 상황에 더욱 민감하게 반응하도록 합니다. 후자는 소수의 긴 응답으로 인해 롤아웃 비용이 증가하는 것을 방지합니다. 두 가지 교사-학생 모델 쌍에서 Adaptive FastOPD는 가장 높은 평균 성능을 달성했으며, OPD 15K와 비교하여 학습 시간을 49.1% ~ 71.2% 단축했습니다. 또한 다양한 하이퍼파라미터 설정에서도 안정적인 성능을 유지합니다.

Original Abstract

On-policy distillation (OPD) provides dense teacher supervision along student-generated trajectories, but its online rollout process incurs substantial computational cost, particularly when a few long responses delay batch completion. Existing acceleration methods typically control rollout length using fixed budgets or absolute teacher--student agreement thresholds, which may not reflect learning progress across different models and training stages. We propose Adaptive FastOPD, a progress-aware strategy that expands the rollout horizon only when learning near the current boundary region has plateaued and the current horizon is sufficiently utilized. The former is determined from four teacher--student signals measured relative to their values upon entering each horizon, making expansion responsive to stage-specific progress rather than a predefined step interval or an absolute threshold on the raw agreement signals, while the latter prevents a small number of long responses from triggering increases in rollout cost. Across two teacher--student pairs, Adaptive FastOPD achieves the highest average performance while reducing training time by 49.1--71.2\% relative to OPD 15K, and remains robust across a range of hyperparameter settings.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!