신뢰 영역 정책 증류
Trust Region Policy Distillation
큰 목표를 한 번에 달성하기 어렵습니다. 대신 작은 단계로 나누는 것이 더 현명합니다. 본 논문에서는 신뢰 영역 정책 증류(Trust Region Policy Distillation, TOP-D)를 제시합니다. TOP-D는 불안정하고 분산이 큰 온라인 정책 증류(On-Policy Distillation, OPD) 방식을 동적으로 근접한 교사 네트워크를 구성하여 안정적인 학습 패러다임으로 전환합니다. 이론적으로, 우리는 TOP-D가 본질적으로 기울기 변동을 제어한다는 것을 입증하는 엄격한 프레임워크를 제시합니다. 또한 전체 학습 역학의 신뢰성과 안정성을 수학적으로 공식화하기 위해, 전역 수렴 분석과 단조적인 성능 향상 경계를 함께 제공합니다. 실험 결과, TOP-D는 수학적 추론 작업에서 학습 안정성, 샘플 효율성 및 최종 성능을 크게 향상시킵니다. 더욱 중요한 점은 TOP-D가 추가적인 계산 오버헤드를 발생시키지 않아, 기존의 OPD 패러다임에 대한 유망한 대안으로 자리매김할 수 있습니다.
Big goals are hard to achieve all at once; breaking them into small steps is wiser. We present Trust Region Policy Distillation (TOP-D), which transforms the notoriously unstable, high-variance On-Policy Distillation (OPD) into a stable training paradigm by dynamically constructing a proximal teacher. Theoretically, we establish a rigorous framework demonstrating that TOP-D inherently controls gradient variance. By providing a formal global convergence analysis alongside a monotonic improvement bound, we mathematically formalize the reliability and stability of the overall training dynamics. Empirically, TOP-D dramatically enhances training stability, sample efficiency, and final performance on mathematical reasoning tasks. More importantly, TOP-D introduces zero additional computational overhead, positioning itself as a promising alternative to the well-established OPD paradigm.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.