선택하고 개선하라: 추론을 위한 추가 학습의 작동 원리 이해
Select and Improve: Understanding the Mechanics of Post-Training for Reasoning
강화 학습은 추론 및 코딩 모델 교육에서 핵심적인 요소로 빠르게 부상했지만, 여전히 메커니즘적인 관점에서 제대로 이해되지 못하고 있습니다. 본 연구는 강화 학습을 통한 추가 학습 과정에서 어떤 방식으로, 그리고 어떠한 근본적인 과정을 통해 능력이 획득되거나 향상되는지를 탐구합니다. Qwen-2.5-1.5B 모델을 대상으로 한 통제된 수학 추론 실험 분석 결과, 전략 선택과 전략 개선이라는 두 가지 핵심 메커니즘이 밝혀졌습니다. 본 연구의 결과는 SFT 데이터와 강화 학습 데이터가 이러한 메커니즘을 활성화하는 데 중요한 역할을 한다는 것을 보여주며, 특히 모델에게 다양한 추론 전략에 대한 지침을 제공함으로써 전략 선택을 가능하게 하고, 강화 학습 데이터의 난이도를 높임으로써 전략 개선을 가능하게 한다는 점을 강조합니다. 종합적으로, 본 연구는 강화 학습 교육에 대한 메커니즘적인 통찰력을 제공하며, 추론 능력을 더욱 발전시키기 위한 실질적인 방안을 제시합니다.
Reinforcement learning has rapidly emerged as a key component in the training of reasoning and coding models, yet it remains poorly understood from a mechanistic perspective. We study how and through what underlying processes capabilities are acquired or enhanced via reinforcement learning post-training. Our analysis, based on controlled math reasoning experiments with Qwen-2.5-1.5B, reveals two core mechanisms: strategy selection and strategy improvement. Our results highlight the role of SFT data and reinforcement learning data in activating these mechanisms, in particular showing how supervising the model on diverse reasoning strategies can enable strategy selection and how increasing difficulty in reinforcement learning data can enable strategy improvement. Taken together, our results provide mechanistic insight into RL training and suggest practical interventions to continue scaling reasoning capabilities.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.