2607.26119v1 Jul 28, 2026 cs.AI

추론 성능의 기원 탐구: 강화 학습(RL)과 지도 학습 기반 미세 조정 모델에서 수학 문제 해결을 위한 표현 품질

Probing the Origins of Reasoning Performance: Representational Quality for Mathematical Problem-Solving in RL vs. SFT Fine-Tuned Models

Akshaj Gurugubelli
Akshaj Gurugubelli
Citations: 0
h-index: 0
Antyabha Rahman
Antyabha Rahman
Citations: 0
h-index: 0
Omar Ankit
Omar Ankit
Citations: 0
h-index: 0
Kevin Zhu
Kevin Zhu
Citations: 19
h-index: 3
Aishwarya H. Balwani
Aishwarya H. Balwani
Citations: 196
h-index: 5

강화 학습(RL)으로 학습된 대규모 추론 모델은 지도 학습 기반 미세 조정(SFT) 모델에 비해 수학적 추론 작업에서 더 높은 성능을 보이는 경우가 많습니다. 하지만 이러한 우위의 근본적인 원인은 아직 명확하지 않습니다. 본 연구에서는 RL 모델이 뛰어난 성능을 발휘하는 데 기여하는 내부 표현 방식의 차이점이 무엇인지 질문합니다. 저희는 두 가지 상호 보완적인 증거를 제시합니다. 첫째, 각 레이어의 숨겨진 상태를 기반으로 학습된 선형 탐침 분석 결과, RL 모델은 SFT 모델보다 정답 예측 정확도가 더 높으며, 이는 더욱 선형적으로 분리 가능하고 구조화된 표현을 나타냅니다. 둘째, 평균적인 중요도 제거(ablation) 연구에서는 RL 모델이 깊은 레이어가 점진적으로 더 중요한 역할을 하는 계층적 아키텍처를 개발하는 반면, SFT 모델은 모든 레이어에 걸쳐 중요도를 균등하게 분배한다는 것을 보여줍니다. 이러한 결과들을 종합하면, RL 학습은 모델이 추론 문제를 표현하고 처리하는 방식을 근본적으로 재구성한다는 것을 알 수 있습니다. 또한, 문제에 대한 반복적인 샘플링을 통해 토큰 수의 변동성을 분석하여 적응적 계산 할당량을 평가했습니다. 일부 RL 기반 모델에서 SFT 기반 모델보다 더 높은 변동성이 관찰되었지만, 다른 모델에서는 일관성이 강한 것으로 나타났습니다. 이는 토큰 할당량이 RL과 SFT 자체보다는 전체 학습 파이프라인에 더 큰 영향을 받는다는 것을 시사합니다. 저희는 이러한 토큰 할당량의 변동성이 가능성 있는 정책의 범위를 보여주며, 어떤 모델은 안정적인 정책을 보이는 반면 어떤 모델은 불확실하거나 식별 불가능한 해결 방식을 나타낼 수 있음을 밝히고자 합니다.

Original Abstract

Large reasoning models trained via reinforcement learning (RL) have been increasingly shown to outperform their supervised fine-tuned (SFT) counterparts on mathematical reasoning tasks; Yet the mechanistic basis for this advantage remains unclear. We therefore ask, what internal representational differences enable RL models' superior performance? Our work presents two converging lines of evidence: First, linear probes trained on layer-wise hidden states reveal that RL models tend to achieve higher accuracy in predicting answer correctness compared to SFT models, indicating more linearly separable and structured representations. Second, mean ablation studies show that RL models develop a hierarchical architecture where deeper layers become progressively more critical, whereas SFT models distribute importance uniformly across layers. Together, these findings demonstrate that RL training fundamentally restructures how models represent and process reasoning problems. Finally, we analyze token-count variability under repeated sampling across problems to assess adaptive compute allocation. While we observe higher variability in some RL-tuned models than in their SFT counterparts, we see strong consistency in others, suggesting that token allocation may depend more on the overall training pipeline than on RL versus SFT alone. We believe this token-allocation variability reveals the spread of plausible on-policy reasoning, highlighting which models exhibit stable policies versus those that are under-determined, potentially non-identifiable solution behaviour.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!