2607.28582v1 Jul 30, 2026 cs.LG

β-OPSD: 정책 최적화를 통해 도출하고, 자기 증류를 통해 학습

$β$-OPSD: Deriving with Policy Optimization, Training with Self-Distillation

Minghui Liu
Minghui Liu
Citations: 8
h-index: 2
Jiawei Xu
Jiawei Xu
Citations: 10
h-index: 1
Juzheng Zhang
Juzheng Zhang
University of Maryland
Citations: 80
h-index: 4
Tom Goldstein
Tom Goldstein
Citations: 82
h-index: 5
Furong Huang
Furong Huang
Citations: 4
h-index: 1

온라인 자기 증류(On-policy Self-Distillation, OPSD)는 추론 언어 모델을 개선하는 데 유망한 접근 방식이지만, 실제 적용 시에는 불안정성이 나타나는 경우가 많습니다. 우리는 이러한 어려움의 근본적인 원인을 구조적 문제에서 찾았습니다. 기본적인 OPSD는 실제로 $β=1$ 값을 갖는 더 넓은 정책 최적화 방법군의 일원이며, 여기서 $β$는 학생 모델을 참조 정책에 고정시키는 KL 페널티를 조절하는 역할을 합니다. 이러한 등가성은 $β$ 값을 단순히 고정된 값으로 취급하던 것을 제어 가능한 정규화 매개변수로 변화시키며, 이는 참조 정책과의 근접성과 우선적으로 제공되는 가이드라인 사이의 균형을 맞출 수 있는 더 일반적인 방법을 제시합니다. 우리는 $β$-OPSD를 소개하고, 그 최적 정책을 참조 정책과 우선 가이드라인 모델 사이의 기하학적 보간으로 정의했습니다. 그러나 이 목표를 강화 학습으로 직접 최적화하는 것은 비용이 많이 들고 분산이 높습니다. 우리는 직접적인 RL 목표 대신, 그 폐쇄형 해를 증류 대상(distillation target)으로 활용합니다. 각 $β$ 값은 참조 모델과 우선 가이드라인 모델 사이의 경로에서 특정 대상을 선택하며, 이를 토큰 수준 로짓 혼합을 통해 효율적으로 구현합니다. 이러한 방식으로 저렴한 증류가 비용이 많이 드는 정책 최적화의 해를 근사합니다. 또한, 미래 예측(return-to-go) 기반 신용 할당은 토큰 업데이트를 시퀀스 레벨 목표와 일치시키면서 OPSD의 단순성을 유지합니다. 수학적 추론 벤치마크에 대한 실험 결과, $β$-OPSD는 기본적인 OPSD보다 일관되게 우수한 성능을 보이며, 최적화 안정성과 후속 추론 성능을 향상시킵니다. 우리의 연구 결과는 자기 증류에서 정책 최적화로, 그리고 다시 자기 증류로의 체계적인 접근 방식을 제시하며, OPSD가 실제로 유용하게 사용될 수 있도록 하는 효율성을 유지합니다.

Original Abstract

On-policy self-distillation (OPSD) is a promising approach to improve reasoning language models, but it remains brittle in practice: making it work reliably often requires substantial engineering effort. We identify a structural source of this difficulty: vanilla OPSD is precisely the $β=1$ member of a broader policy-optimization family, where $β$ weights the KL penalty anchoring the student to a reference policy. This equivalence turns $β$ from an implicit value fixed at one into a controllable regularization parameter, yielding a more general formulation that trades off proximity to a reference policy against privileged teacher guidance. We introduce $β$-OPSD and derive its optimal policy as a geometric interpolation between the reference policy and the privileged teacher. Directly optimizing this objective with reinforcement learning, however, would be costly and high-variance. Rather than optimize the RL objective directly, we turn its closed-form solution into a distillation target. Each value of $β$ selects a target along the reference-to-teacher path, which we implement efficiently by mixing their token-level logits. In this way, inexpensive distillation approximates the solution of expensive policy optimization. Return-to-go credit assignment further aligns token updates with the sequence-level objective while retaining the simplicity of OPSD. Experiments on mathematical reasoning benchmarks show that $β$-OPSD consistently outperforms vanilla OPSD, improving optimization stability and downstream reasoning performance. Our results provide a principled route from self-distillation to policy optimization and back without sacrificing the efficiency that makes OPSD practical.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!