추론 깊이와 환경 복잡성: 논리적 추론 작업에서의 RLVR 데이터 할당에 대한 통제된 연구
Reasoning Depth and Environment Complexity: A Controlled Study of RLVR Data Allocation across Logical Reasoning Tasks
검증 가능한 보상을 사용하는 강화 학습(RLVR)은 사후 학습 추론 모델의 핵심 요소가 되었지만, 기존 연구의 주요 한계점은 추론 공간에 대한 지나치게 제한적인 관점입니다. 즉, 난이도는 단순히 추론 깊이만으로 정의되고, 보상은 주로 전방향 연역적 상태 추적에 집중됩니다. 본 연구에서는 추론 공간을 두 가지 차원으로 특징짓습니다. 첫째, 추론 깊이 외에도 환경 복잡성을 고려하며, 모델이 주의를 분산시키는 요소와 상호 작용 구조 속에서 올바른 경로를 식별해야 하는 경우를 포함합니다. 둘째, 실제 세계 추론에 핵심적인 네 가지 능력을 다룹니다: 연역적 상태 추적, 숨겨진 사건 또는 사실의 귀납적 회복, 유도 규칙 생성, 그리고 유사성 전이. 이러한 요인들을 분리하기 위해, 통제된 사전 및 사후 학습 분포를 갖는 합성 지식 그래프 환경을 구축했으며, 각 샘플은 깊이, 복잡성 및 작업 유형에 따라 다양하게 구성됩니다. 연구 결과 다음과 같은 세 가지 사실이 밝혀졌습니다: 깊이와 복잡성을 함께 고려한 방법이 단일 축 기반 방법에 비해 우수한 성능을 보입니다; 추론 유형별로 반응이 균일하지 않으며, 특히 귀납적 추론은 RL에서 다루는 영역 밖에서 성능 저하를 나타내고, 작업 간의 상관관계는 연역-귀납 및 유도-유사성 쌍으로 묶이는 경향이 있습니다; 마지막으로, 고정된 예산 하에서는 단계별 학습 방식보다 균일한 혼합 방식이 더 효과적입니다. 또한, 최근 출시된 상용 모델에서도 동일한 연역 우세-귀납 약점 현상이 나타나는 것을 확인했으며, 이는 이러한 격차가 본 연구의 통제된 환경 설정으로 인해 발생하는 것이 아님을 시사합니다.
Reinforcement learning with verifiable rewards (RLVR) has become central to post-training reasoning models, yet a key limitation of existing studies is their narrow view of the reasoning space: difficulty is treated as reasoning depth alone, and reward is concentrated on forward deductive state tracking. We instead characterize the reasoning space along two dimensions. Difficulty. Beyond reasoning depth, we study environment complexity, where models must identify the correct path amid distractors and interacting structures. Rewarded reasoning form. We consider four abilities core to real-world reasoning: deductive state tracking, abductive recovery of hidden events or facts, inductive rule induction, and analogical transfer. To disentangle these factors, we construct a synthetic knowledge-graph environment with controlled pre- and post-training distributions, where each instance varies along depth, complexity, and task family. Three findings emerge: joint depth-complexity coverage outperforms single-axis recipes; reasoning families respond non-uniformly, with abductive reasoning degrading outside the RL-covered region and task correlations clustering into deductive-abductive and inductive-analogy pairs; and uniform mixing outperforms staged curricula under a fixed budget. We also find that recent off-the-shelf models exhibit the same deductive-over-abductive asymmetry, suggesting that this gap is not merely an artifact of our controlled setup.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.