2605.28388v1 May 27, 2026 cs.AI

LLM의 강화 학습 기반 검증 가능한 보상(RLVR)에서 샘플 난이도가 갖는 역할에 대한 메커니즘적 해석

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs

Jiajun Zhang
Jiajun Zhang
Citations: 1,391
h-index: 7
Zheng Wang
Zheng Wang
Citations: 17
h-index: 1
Weiwei Xing
Weiwei Xing
Citations: 12
h-index: 2
Zhanxing Zhu
Zhanxing Zhu
Citations: 50
h-index: 4
Yue Cheng
Yue Cheng
Citations: 14
h-index: 2
Xiaohui Gao
Xiaohui Gao
Citations: 6
h-index: 2

검증 가능한 보상을 이용한 강화 학습 (RLVR)은 특히 수학 및 프로그래밍 분야에서 대규모 언어 모델 (LLM)의 추론 성능을 현저하게 향상시키는 것으로 경험적으로 입증되었습니다. 그러나 RLVR에서 샘플 난이도가 갖는 메커니즘적 역할에 대한 이해는 아직 부족합니다. 본 논문에서는 난이도별 분석 및 단일 샘플 분석 관점에서 RLVR을 연구합니다. 우리는 샘플 난이도가 RLVR에 미치는 영향이 단조적이지 않다는 것을 발견했습니다. 쉬운 문제와 중간 난이도의 문제가 가장 강력하고 안정적인 추론 성능 향상을 가져오는 반면, 지나치게 어려운 문제는 종종 약한 학습 신호를 제공하며, 답의 반복 또는 필요한 계산을 건너뛰는 것과 같은 퇴화된 행동을 유발하며, 궁극적으로 모델의 기존 능력을 저하시킬 수 있습니다. 응답 외에도, 우리는 Temporal Sparse Autoencoders (T-SAE)를 사용하여 모델의 내부 특징 동역학을 추가로 분석했습니다. 쉬운 문제는 주로 직접 답변 및 기본 계산 관련 특징을 강화하고 심층적 추론 관련 특징을 억제하는 반면, 어려운 문제는 추론 관련 특징을 활성화하지만 성공적인 경로가 샘플링될 때만 유용합니다. 중간 난이도의 문제는 계산과 다단계 추론 모두를 강화하여 더욱 균형 잡힌 신호를 제공합니다. 이러한 연구 결과를 바탕으로, 우리는 역추론 재구성 및 T-SAE 기반 학습 신호를 사용하여 보상 밀도와 RLVR 과정에서의 공헌도 할당을 개선하는 어려운 샘플 활용을 위한 난이도 적응 전략을 제안합니다. 전반적으로, 본 논문의 결과는 샘플 난이도가 RLVR의 최적화 동역학과 표현 진화를 규제하는 핵심 요소임을 보여줍니다.

Original Abstract

Reinforcement Learning with Verifiable Reward (RLVR) is empirically shown to notably enhance the reasoning performance of large language models (LLMs), particularly in mathematics and programming. However, the mechanistic role of Sample Difficulty in RLVR remains poorly understood. In this paper, we investigate RLVR through the lens of difficulty-wise and one-sample analysis. We find that sample difficulty has a non-monotonic effect on RLVR: easy and medium-difficulty problems yield the strongest and most stable reasoning improvements, whereas overly hard problems often provide weak learning signals, induce degenerate behaviors such as answer repetition or skipping necessary computation, and can ultimately degrade the model's pre-existing capabilities. Beyond the obverse of response, we further analyze the model's internal feature dynamics using Temporal Sparse Autoencoders (T-SAE). Easy problems mainly reinforce direct-answer and basic-computation features while suppressing deliberative-reasoning features; hard problems activate reasoning-related features but become useful only when successful trajectories are sampled; medium-difficulty problems provide a more balanced signal, strengthening both computation and multi-step reasoning features. Motivated by these findings, we propose difficulty-adaptive strategies for hard-sample utilization, using backward-reasoning reformulation and T-SAE-guided training signals to improve reward density and credit assignment during RLVR. Overall, our results identify sample difficulty as a key factor governing both the optimization dynamics and representation evolution of RLVR.

1 Citations
0 Influential
3.5 Altmetric
18.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!