2604.05134v1 Apr 06, 2026 cs.LG

체스를 통한 추론: 데이터 기반의 미세 조정과 강화 학습을 통해 추론이 어떻게 발전하는가

Reasoning Through Chess: How Reasoning Evolves from Data Through Fine-Tuning and Reinforcement Learning

Lucas Dionisopoulos
Lucas Dionisopoulos
Citations: 0
h-index: 0
Nicklas Majamaki
Nicklas Majamaki
Citations: 14
h-index: 1
Prithviraj Ammanabrolu
Prithviraj Ammanabrolu
Citations: 3,194
h-index: 25

언어 모델이 본질적으로 어려움을 겪는 작업에서 어떻게 추론 능력을 갖추게 할 수 있을까요? 본 연구는 이론적으로 설계된 데이터셋이 언어 모델의 체스 수행 능력에 미치는 영향을 분석하여, 언어 모델의 추론이 지도 학습 기반 미세 조정(SFT)에서 강화 학습(RL)으로 어떻게 발전하는지 연구합니다. 모델이 최적의 수를 직접 예측하도록 미세 조정하는 것이 효과적인 강화 학습과 최고의 성능을 가져온다는 것을 확인했지만, 강화 학습 단계에서는 선택된 수와 일치하지 않는 부실한 추론이 나타났습니다. 반면에, 다중 수의 경로를 사용하여 학습하면, 유사한 성능을 보이면서도 더 신뢰할 수 있는 추론과 안정적인 강화 학습을 얻을 수 있습니다. 또한, 강화 학습은 수의 품질 분포를 긍정적으로 변화시키고, 환각 현상 발생률을 감소시키는 효과도 있습니다. 마지막으로, 평가 성능, 환각 현상 발생률, 추론 품질 등 다양한 SFT 중간 결과 지표가 강화 학습 이후의 모델 성능을 예측하는 데 유용하다는 것을 확인했습니다. 본 연구에서는 학습 데이터, 평가 결과, 코드와 함께 중간 결과 모델 및 최종 모델을 공개하며, 이를 통해 70억 파라미터의 모델로 선도적인 오픈 소스 추론 모델을 능가하는 성과를 달성했습니다.

Original Abstract

How can you get a language model to reason in a task it natively struggles with? We study how reasoning evolves in a language model -- from supervised fine-tuning (SFT) to reinforcement learning (RL) -- by analyzing how a set of theoretically-inspired datasets impacts language model performance in chess. We find that fine-tuning a model to directly predict the best move leads to effective RL and the strongest downstream performance -- however, the RL step elicits unfaithful reasoning (reasoning inconsistent with the chosen move). Alternatively, training on multi-move trajectories yields comparable downstream performance with faithful reasoning and more stable RL. We show that RL induces a substantial positive shift in the distribution of move quality and reduces hallucination rates as a side effect. Finally, we find several SFT-checkpoint metrics -- metrics spanning evaluation performance, hallucination rates, and reasoning quality -- to be predictive of post-RL model performance. We release checkpoints and final models as well as training data, evaluations, and code which allowed us to surpass leading open-source reasoning models in chess with a 7B-parameter model.

0 Citations
0 Influential
12.5 Altmetric
62.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!