DRIFT: 리듬 기반 탐색과 성공 버퍼 학습을 통한 자체 증류의 난이도 기반 라우팅
DRIFT: Difficulty Routing Self-DIstillation with Rhythm-Gated Exploration and Success BuFfer Training
외부 전문가 감독 없이 대규모 언어 모델이 안정적인 자기 개선을 달성하는 것은 복잡한 추론 작업에서 여전히 중요한 과제입니다. 기존의 자체 증류 및 강화 학습 방법은 문제 수준에서의 학습 진행 상황을 명시적으로 추적하고 그에 따라 최적화 전략을 조정할 수 있는 메커니즘이 부족합니다. 결과적으로, 훈련 과정에서 쉬운 문제에 대해 과도하게 최적화될 수 있으며, 어려운 문제로부터는 미약한 지침을 받을 수 있고, 경계 사례를 충분히 탐색하지 못할 수 있습니다. 이러한 문제를 해결하기 위해, 우리는 대규모 언어 모델을 위한 온라인 자체 진화 정책 최적화 프레임워크인 DRIFT를 제안합니다. DRIFT는 난이도 기반 라우팅과 리듬 게이팅의 결합된 사용을 통해 모델의 자기 개선 과정을 조절합니다. 전자는 문제 수준에서 모델의 학습 상태를 파악하고 자체 증류 및 강화 학습 신호를 동적으로 할당하며, 후자는 토큰 수준에서 정책 업데이트를 개선하여 중요한 추론 위치에 대한 탐색을 집중시킵니다. 또한, 성공 버퍼와 두 단계의 교육 과정 학습 전략을 추가적으로 통합하여 DRIFT는 고품질의 과거 경험을 보존하면서 모델이 안정적인 정책 진화로 이어지도록 점진적으로 안내합니다. 다섯 가지 벤치마크 및 세 가지 모델 크기로 평가된 결과, DRIFT는 GRPO와 SDPO 모두를 능가하는 최상의 성능을 보였습니다. 다섯 가지 벤치마크의 평균 점수에서 DRIFT는 79.5%를 달성하여 GRPO보다 9.5%, SDPO보다 7.5% 더 높은 수치를 기록하며 새로운 최고 수준의 결과를 보여주었습니다. 특히, ToolUse 작업에서 DRIFT는 정확도가 79.2%로 GRPO보다 13.5%, SDPO보다 10.7% 향상되어 새로운 최고 수준을 설정했으며, 동시에 다른 모든 동시대 방법보다 훨씬 우수한 성능을 보였습니다.
Enabling large language models to achieve stable self-improvement without external expert supervision remains a central challenge in complex reasoning tasks. Existing self-distillation and reinforcement learning methods lack explicit mechanisms for tracking problem-level learning progress and adapting optimization strategies accordingly. Consequently, training may over-optimize easy problems, receive weak supervision from hard problems, and fail to sufficiently explore borderline cases. To resolve these issues, we propose DRIFT, an online self-evolution policy optimization framework for large language models. DRIFT regulates the model's self-improvement process through the joint use of Difficulty Routing and Rhythm Gating. The former identifies the model's learning state at the problem level and dynamically allocates self-distillation and reinforcement learning signals, while the latter refines policy updates at the token level, concentrating exploration on critical reasoning positions. By further incorporating a success buffer and a two-stage curriculum learning strategy, DRIFT preserves high-quality historical experience while progressively guiding the model from reliable behavior acquisition toward stable policy evolution. Evaluated across five benchmarks and three model scales, DRIFT surpasses the peak performance of both GRPO and SDPO across all evaluated metrics. On the average score over the five benchmarks, DRIFT achieves 79.5$\%$, outperforming GRPO by 9.5$\%$ and SDPO by 7.5$\%$, establishing a new state-of-the-art result. Notably, on ToolUse, DRIFT reaches an accuracy of 79.2$\%$, improving over GRPO by 13.5$\%$ and SDPO by 10.7$\%$, setting a new state-of-the-art and substantially outperforming all concurrent methods.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.