2607.04751v1 Jul 06, 2026 cs.LG

신뢰 영역 정책 증류

Trust Region Policy Distillation

Zeke Xie
Zeke Xie
Citations: 371
h-index: 8
Zhengpeng Xie
Zhengpeng Xie
Citations: 31
h-index: 2
L. Zhang
L. Zhang
Citations: 3,688
h-index: 13
Mao Yang
Mao Yang
Citations: 1,035
h-index: 12

큰 목표를 한 번에 달성하기 어렵습니다. 대신 작은 단계로 나누는 것이 더 현명합니다. 본 논문에서는 신뢰 영역 정책 증류(Trust Region Policy Distillation, TOP-D)를 제시합니다. TOP-D는 불안정하고 분산이 큰 온라인 정책 증류(On-Policy Distillation, OPD) 방식을 동적으로 근접한 교사 네트워크를 구성하여 안정적인 학습 패러다임으로 전환합니다. 이론적으로, 우리는 TOP-D가 본질적으로 기울기 변동을 제어한다는 것을 입증하는 엄격한 프레임워크를 제시합니다. 또한 전체 학습 역학의 신뢰성과 안정성을 수학적으로 공식화하기 위해, 전역 수렴 분석과 단조적인 성능 향상 경계를 함께 제공합니다. 실험 결과, TOP-D는 수학적 추론 작업에서 학습 안정성, 샘플 효율성 및 최종 성능을 크게 향상시킵니다. 더욱 중요한 점은 TOP-D가 추가적인 계산 오버헤드를 발생시키지 않아, 기존의 OPD 패러다임에 대한 유망한 대안으로 자리매김할 수 있습니다.

Original Abstract

Big goals are hard to achieve all at once; breaking them into small steps is wiser. We present Trust Region Policy Distillation (TOP-D), which transforms the notoriously unstable, high-variance On-Policy Distillation (OPD) into a stable training paradigm by dynamically constructing a proximal teacher. Theoretically, we establish a rigorous framework demonstrating that TOP-D inherently controls gradient variance. By providing a formal global convergence analysis alongside a monotonic improvement bound, we mathematically formalize the reliability and stability of the overall training dynamics. Empirically, TOP-D dramatically enhances training stability, sample efficiency, and final performance on mathematical reasoning tasks. More importantly, TOP-D introduces zero additional computational overhead, positioning itself as a promising alternative to the well-established OPD paradigm.

0 Citations
0 Influential
6.5 Altmetric
32.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!