2606.18831v1 Jun 17, 2026 cs.CL

보상을 넘어: 장문 컨텍스트 강화 학습을 위한 데이터 레시피

Beyond Reward Engineering: A Data Recipe for Long-Context Reinforcement Learning

Xiaorong Wang
Xiaorong Wang
Citations: 50
h-index: 3
Xiaoyue Xu
Xiaoyue Xu
Citations: 24
h-index: 2
Chaojun Xiao
Chaojun Xiao
Citations: 3,593
h-index: 24
Xu Han
Xu Han
Citations: 106
h-index: 3
Si Zhang
Si Zhang
Citations: 79
h-index: 3

장문 컨텍스트 추론은 특히 자율 에이전트로 활용될 때, 대규모 언어 모델에게 필수적인 능력입니다. 최근 강화 학습(RL)은 이러한 능력을 향상시키는 주요 패러다임으로 떠올랐지만, 기존 연구는 주로 보상 설계에 집중하는 반면 다양한 학습 데이터는 여전히 부족합니다. 본 논문에서는 데이터 중심의 관점에서 이 문제를 재검토하고, 간단하면서도 효과적인 데이터 레시피와 최소한의 결과 기반 GRPO 설정을 결합하는 것만으로도 장문 컨텍스트 추론 능력을 크게 향상시킬 수 있음을 보여줍니다. 저희의 레시피는 검색, 다중 증거 종합, 그리고 추론이라는 세 가지 상호 보완적인 작업 유형을 대상으로 하며, 이를 위해 약 14,000개의 예제로 구성된 8개의 데이터셋을 구축하고 관리했습니다. Qwen3-4B/8B/30B-A3B 모델 세 개에 대한 실험 결과, 저희의 레시피를 사용했을 때 평균적으로 7개 이상의 장문 컨텍스트 벤치마크에서 +7.2/+3.2/+6.4점의 성능 향상을 보였으며, 이는 기존 강화 학습 데이터셋을 능가하는 결과입니다. 또한, 이러한 성능 향상이 에이전트 기반 작업에도 적용될 수 있음을 보여주며, 저희의 데이터 레시피를 사용하여 에이전트에 최적화된 모델을 추가적으로 훈련시킨 결과, GAIA에서 +4.8점, BrowseComp에서 +7.0점의 성능 향상을 달성했습니다. 본 연구에서 사용한 데이터셋은 향후 연구에 활용될 수 있도록 공개할 예정입니다.

Original Abstract

Long-context reasoning is an essential capability for large language models, particularly when they are deployed as autonomous agents that must reason over lengthy trajectories. Reinforcement learning (RL) has recently emerged as a dominant paradigm for improving this ability, yet existing work largely focuses on reward engineering while diverse training data remains scarce. We revisit this problem from a data-centric perspective and show that a simple yet effective data recipe alone, paired with a minimal outcome-based GRPO setup, suffices to substantially improve long-context reasoning. Our recipe targets three complementary task families -- retrieval, multi-evidence synthesis, and reasoning -- for which we construct and curate eight datasets totaling ~14K examples. Experiments on three models (Qwen3-4B/8B/30B-A3B) yield average gains of +7.2/+3.2/+6.4 points across seven long-context benchmarks, surpassing prior RL training sets. We further demonstrate that these gains transfer to agentic tasks, where continuing RL training on an agent-tuned model with our data recipe improves GAIA by +4.8 and BrowseComp by +7.0 points. We will release our datasets to facilitate future research.

0 Citations
0 Influential
12 Altmetric
60.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!