2606.31575v1 Jun 30, 2026 cs.AI

어떤 토큰이 중요한가? 상대적 놀람 지수(Relative Surprisal Index)를 활용한 강화학습 기반 검증 가능한 보상(RLVR)에서의 적응적 토큰 선택

Which Tokens Matter? Adaptive Token Selection for RLVR with the Relative Surprisal Index

Yanzhao Zheng
Yanzhao Zheng
Citations: 84
h-index: 4
Baohua Dong
Baohua Dong
Citations: 66
h-index: 4
Hangcheng Zhu
Hangcheng Zhu
Citations: 44
h-index: 3
Outongyi Lv
Outongyi Lv
Citations: 56
h-index: 3
Yuanwei Zhang
Yuanwei Zhang
Citations: 0
h-index: 0
Zhenghao Huang
Zhenghao Huang
Citations: 0
h-index: 0
Yingda Chen
Yingda Chen
Citations: 460
h-index: 6
Xingjun Wang
Xingjun Wang
Citations: 0
h-index: 0

강화 학습(Reinforcement Learning, RL)은 대규모 언어 모델(Large Language Models, LLMs)을 단순 모방 훈련에서 벗어나 더욱 강력한 추론 능력을 갖도록 발전시키는 데 중요한 도구로 자리 잡았습니다. 기존 접근 방식 중, 검증 가능한 보상 기반 강화 학습(Reinforcement Learning with Verifiable Rewards, RLVR)은 LLM의 추론 능력 향상을 위한 핵심 패러다임으로 부상했습니다. 하지만 최근 연구들은 서로 다른 관점을 제시하며, 일부는 훈련 시 높은 엔트로피를 가진 토큰 위치를 우선시해야 한다고 주장하는 반면, 다른 일부는 낮은 확률을 가진 토큰이 기울기 업데이트를 지배하도록 허용해서는 안 된다고 경고합니다. 주목할 점은, 일반적으로 높은 엔트로피를 가진 토큰은 낮은 확률과 상관관계가 있지만, 두 가지 접근 방식 모두 경험적으로 상당한 성능 향상을 가져온다는 것입니다. 본 연구에서는 샘플링된 토큰의 확률 또는 엔트로피를 개별적으로 평가하는 것만으로는 정책 최적화 역학을 충분히 설명할 수 없다고 주장합니다. 이러한 문제점을 해결하기 위해, 본 연구는 토큰의 엔트로피와 선택된 토큰의 확률을 자연스럽게 결합하는 정보 이론 기반 지표인 상대적 놀람 지수(Relative Surprisal Index, RSI)를 제안합니다. 우리는 비교적 완만한 조건 하에서 RSI가 선택된 로짓에 대한 예측 엔트로피 및 일차 변동의 로짓 기울기 정규화 비율과 관련이 있음을 보여줍니다. RSI를 기반으로, 본 연구는 안정적인 RSI 구간 내에 있는 토큰을 유지하는 엔트로피 적응형 토큰 필터링 방법인 RSI Selection (RSI-S)을 제안합니다. RSI-S는 기존의 상반된 패러다임을 성공적으로 조화시키고, 중복되는 낮은 놀람 지수의 토큰과 불안정한 높은 놀람 지수의 꼬리 부분에 있는 토큰을 모두 제거합니다. 다양한 모델 크기(Qwen2.5-1.5B, 3B, 및 7B)에서 AIME 및 AMC 벤치마크를 사용하여 수행된 실험 결과, RSI-S는 GRPO보다 avg@32 정확도가 2~3% 향상되는 것을 확인했습니다. 전반적으로, RSI는 RLVR 개선을 위한 유망한 관점을 제시합니다.

Original Abstract

Reinforcement learning (RL) has become a powerful tool for propelling Large Language Models (LLMs) beyond imitation-based training towards more robust reasoning capabilities. Among existing approaches, RL with Verifiable Rewards (RLVR) has emerged as a pivotal paradigm for advancing LLM reasoning. Despite its empirical success, recent studies have offered different insights. One line of inquiry advocates prioritizing high-entropy token positions during training, while another perspective cautions against allowing low-probability tokens to dominate gradient updates. Notably, although high-entropy tokens are usually correlated with low probability, both paradigms empirically yield substantial performance gains. In this work, we argue that evaluating sampled-token probability or entropy in isolation is insufficient to capture the policy optimization dynamics. To resolve this tension, we introduce the Relative Surprisal Index (RSI), a principled, information-theoretic metric that naturally couples the token's entropy with the probability of the selected token. We show that, under mild conditions, RSI is related to the local ratio between the first-order variations of the logit-gradient norm and predictive entropy under a selected-logit perturbation. Building on RSI, we propose RSI Selection (RSI-S), an entropy-adaptive token filtering method that retains tokens within a stable RSI interval. RSI-S successfully reconciles previous contradictory paradigms and filters out both redundant low-surprisal tokens and unstable high-surprisal tail tokens. Empirical evaluations show that RSI-S achieves higher avg@32 accuracy across different model scales (Qwen2.5-1.5B, 3B, and 7B) on AIME and AMC benchmarks: RSI-S improves avg@32 accuracy by 2--3 percentage points over GRPO. Overall, RSI offers a promising perspective for RLVR improvement.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!