시작 지점의 다양성: RLVR을 위한 낮은 부하, 높은 효과의 초기 토큰 다양화
Where Rollouts Begin: Low-Load, High-Leverage First-Token Diversification for RLVR
검증 가능한 보상을 활용한 강화 학습(RLVR)은 레이블이 없는 추론 경로를 사용하여 추론 모델을 훈련하며, 정책이 다양한 추론 경로를 경험할 수 있도록 그룹으로 묶인 실행(rollout)을 사용하고, 검증기를 통해 이를 평가합니다. 따라서 실행의 다양성은 RLVR에서 중요한 제약 요인이 되었으며, 대부분 기존 방법은 온도 조절, 접두사 추가 또는 실행 선택 조정 등을 통해 탐색 범위를 넓히는 방식을 사용합니다. 본 연구에서는 구조적으로 구별되지만 간과되었던 다양성 확대 지점인 추론 마커 바로 다음 토큰의 중요성을 강조합니다. 정책의 첫 번째 토큰 분포는 뚜렷한 피크를 가지면서도 정확성과 독립적인 현상을 보이며, 이 첫 번째 토큰 위치는 실행 그룹이 커버하는 영역을 넓히면서 동시에 정확성 신호를 변경하지 않습니다. 우리는 REFT(Rollout Exploration with First-Token Diversification)라는 가벼운 추가 기능을 RLVR 파이프라인에 도입하여, 정책의 상위 $N$ 후보에서 첫 번째 토큰을 균등하게 샘플링하고 실행을 고르게 분배하며, 나머지 구성 요소는 그대로 유지합니다. REFT는 이렇게 다양화된 실행으로 훈련되었으며, 네 가지 기본 모델(0.5B-7B)과 세 가지 난이도 수준에서 DAPO 및 GRPO 기준 성능보다 Pass@1, Pass@8, Pass@64의 집계 성능을 향상시켰습니다.
Reinforcement Learning with Verifiable Rewards (RLVR) trains reasoning models without labeled trajectories, relying on grouped rollouts to expose the policy to alternative reasoning paths and a verifier to score them. Rollout diversity has accordingly emerged as a central bottleneck in RLVR, with most existing methods broadening exploration through temperature, prefix, or rollout-selection adjustments. We identify a structurally distinguished but overlooked position for broadening this diversity: the first token after the reasoning marker. The policy's first-token distribution exhibits a sharply peaked yet correctness-decoupled phenomenon, and this first token position can broaden the regions a rollout group covers without altering the correctness signal. We introduce REFT (Rollout Exploration with First-Token Diversification), a light addition to the RLVR pipeline that samples first tokens uniformly from the policy's own top-$N$ candidates and allocates rollouts evenly, leaving every other component unchanged. Trained on the resulting diversified rollouts, REFT improves aggregate Pass@1, Pass@8, and Pass@64 over DAPO and GRPO baselines across four base models (0.5B-7B) and three difficulty regimes.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.