2608.03119v1 Aug 04, 2026 cs.AI

정답을 보지 마세요: 결과 가려 그룹 상대 정책 최적화를 통한 레이블 없는 강화 학습 기반 언어 모델

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR

Yidong Chen
Yidong Chen
Citations: 116
h-index: 5
Biao Fu
Biao Fu
Xiamen University
Citations: 201
h-index: 8
Yongshi Ye
Yongshi Ye
Citations: 17
h-index: 3
Xiaodong Shi
Xiaodong Shi
Citations: 56
h-index: 5
Liang Zhang
Liang Zhang
Citations: 84
h-index: 5

검증 가능한 보상을 활용한 강화 학습(RLVR)은 LLM의 추론 능력을 향상시키지만, 일반적으로 정답 데이터에 의존하여 확장성에 제한이 있습니다. 투표 기반의 레이블 없는 RLVR은 정답 데이터를 대신하여 모델 샘플로부터 얻은 답변 수준의 합의를 사용합니다. 그러나 동일한 답변 수준 신호를 보상 추정과 토큰 수준 정책 최적화 모두에 사용할 경우, 모델이 추론 능력을 향상시키는 대신 직접적으로 답변 토큰을 강화하도록 유도하여 성능 저하가 발생할 수 있습니다. 본 연구에서는 RLVR 프레임워크인 OM-GRPO를 제안합니다. OM-GRPO는 보상 추정과 정책 최적화를 분리하며, 답변 범위에 대한 그래디언트를 가리고 소프트한 합의 신호를 통해 답변 수준 보상을 유지함으로써 최적화 압력을 답변 토큰에서 벗어나게 합니다. 또한, 기존 트래젝토리에 대한 저렴한 쌍별 비교를 통해 보상 추정을 개선하는 Contrast-Augmented Reward 기법을 도입했습니다. 다양한 추론 벤치마크와 세 가지 LLM 백본 모델에 대해 OM-GRPO는 기존의 레이블 없는 RLVR 방법보다 우수한 성능을 보이며, 지도 학습 기반의 정답 데이터 활용 방식과 안정적인 최적화를 통해 동등한 수준의 성능을 달성합니다. 특히 Test-Time Training 환경에서 OM-GRPO는 다수 투표 방식에 비해 4.24점 더 높은 성능을 보여줍니다.

Original Abstract

Reinforcement Learning with Verifiable Rewards (RLVR) improves LLM reasoning but typically relies on ground-truth (GT) answers, limiting scalability. Voting-based label-free RLVR replace gold supervision with answer-level consensus from model samples. However, collapse arises when the same answer-level signal is used both to estimate rewards and to drive token-level policy optimization, encouraging the model to directly reinforce answer tokens rather than improve reasoning. We propose OM-GRPO, a label-free RLVR framework that decouples reward estimation from policy optimization. OM-GRPO masks gradients on the answer span while retaining answer-level rewards through a soft consensus signal, shifting optimization pressure away from answer tokens. We further introduce Contrast-Augmented Reward, which refines reward estimation via low-cost pairwise comparisons over existing trajectories without additional rollouts. Across diverse reasoning benchmarks and three LLM backbones, OM-GRPO consistently outperforms existing label-free RLVR methods and matches supervised GT-reward training with stable optimization. This stability is particularly beneficial in the Test-Time Training setting, where OM-GRPO surpasses majority voting by 4.24 points.

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!