사후 훈련 과정에서 지도 학습 미세 조정과 강화 학습의 결합성에 대한 연구
On the Non-decoupling of Supervised Fine-tuning and Reinforcement Learning in Post-training
대규모 언어 모델의 사후 훈련 과정에서 지도 학습 미세 조정(SFT)과 강화 학습(RL)이 일반적으로 함께 사용됩니다. 이 두 방법은 서로 다른 목표를 가지고 있습니다. SFT는 모델 출력과 전문가 답변 간의 교차 엔트로피 손실을 최소화하는 반면, RL은 인간의 선호도 또는 규칙 기반 검증 시스템에서 파생된 보상 신호를 최대화합니다. 현대적인 추론 모델은 SFT와 RL 훈련을 번갈아 사용하는 방식을 널리 채택하고 있습니다. 그러나 이러한 두 방법이 분리될 수 있는지에 대한 이론적 근거는 아직 없습니다. 본 연구에서는 분리가 불가능함을 증명합니다. (1) SFT-then-RL 결합: RL은 SFT 최적 조건에서 SFT 손실을 증가시킵니다. (2) RL-then-SFT 결합: SFT는 RL이 달성하는 보상을 감소시킵니다. Qwen3-0.6B 모델에 대한 실험 결과는 예측된 성능 저하를 확인시켜 주며, SFT와 RL이 사후 훈련 과정에서 이전 성능 손실 없이 분리될 수 없음을 입증합니다.
Post-training of large language models routinely interleaves supervised fine-tuning (SFT) with reinforcement learning (RL). These two methods have different objectives: SFT minimizes the cross-entropy loss between model outputs and expert responses, while RL maximizes reward signals derived from human preferences or rule-based verifiers. Modern reasoning models have widely adopted the practice of alternating SFT and RL training. However, there is no theoretical account of whether they can be decoupled. We prove that decoupling is impossible in either order: (1) SFT-then-RL coupling: RL increases SFT loss under SFT optimality and (2) RL-then-SFT coupling: SFT lowers the reward achieved by RL. Experiments on Qwen3-0.6B confirm the predicted degradation, verifying that SFT and RL cannot be separated without loss of prior performance in the post-training
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.