STARE: 서프라이즈 기반 토큰 레벨 이점 재가중화 기법을 통한 정책 엔트로피 안정성 확보
STARE: Surprisal-Guided Token-Level Advantage Reweighting for Policy Entropy Stability
GRPO와 같은 검증 가능한 보상을 사용하는 강화 학습 알고리즘은 LLM의 복잡한 추론 능력 향상을 위한 중요한 방법으로 떠오르고 있지만, 훈련 과정에서 종종 정책 엔트로피 감소 문제를 겪습니다. 본 연구에서는 GRPO 환경 하에서의 토큰 레벨 엔트로피 변화에 대한 1차 기울기 분석을 수행하고, 토큰 레벨 신용 할당 불일치를 확인했습니다. 각 토큰의 엔트로피 변동은 전체 경로 레벨 이점과 다음 토큰 분포에 대한 민감도 함수의 곱으로 분해되며, 이를 통해 이점-서프라이즈 간의 4분면 구조와 거의 임계적인 특성을 발견했습니다. 이러한 분석을 바탕으로, 본 연구에서는 정책 엔트로피 안정화를 위한 서프라이즈 기반 토큰 레벨 이점 재가중화 기법인 STARE를 제안합니다. STARE는 배치 내에서 서프라이즈 분위수를 활용하여 엔트로피 민감도가 높은 토큰 집합을 식별하고, 선택적으로 해당 토큰들의 효과적인 이점을 재가중하며, 안정적인 엔트로피 조절을 위한 목표 엔트로피 폐루프 게이트를 통합합니다. 1.5B에서 32B까지 다양한 모델 크기와 Short CoT, Long CoT, Multi-Turn Tool Use의 세 가지 작업 유형에 대해 STARE는 수천 단계 동안 안정적인 강화 학습을 유지하면서 정책 엔트로피를 목표 범위 내로 관리합니다. AIME24 및 AIME25 데이터셋에서 STARE는 DAPO 및 기타 경쟁 모델 대비 평균 정확도 4%-8% 향상을 보여주었으며, 반사 토큰과 응답 길이를 함께 증가시켜 지속적인 탐색-활용 균형을 유지하며 강화 학습 잠재력을 더욱 향상시킵니다. 관련 코드는 https://github.com/hp-luo/STARE 에서 확인할 수 있습니다.
Reinforcement Learning with Verifiable Rewards algorithms like GRPO have emerged as the dominant post-training paradigm for complex reasoning in LLMs, yet commonly suffer from policy entropy collapse during training. We conduct a first-order gradient analysis of token-level entropy dynamics under GRPO and identify a token-level credit assignment mismatch: the per-token entropy variation decomposes into the product of the trajectory-level advantage and an entropy sensitivity function over the next-token distribution, yielding an advantage-surprisal four-quadrant structure and a near-criticality property. Motivated by it, we propose STARE (Surprisal-guided Token-level Advantage Reweighting for policy Entropy stability), which identifies entropy-critical token subsets via batch-internal surprisal quantiles, selectively reweights their effective advantages, and incorporates a target-entropy closed-loop gate for stable entropy regulation. Across model scales from 1.5B to 32B and three task families (Short CoT, Long CoT, and Multi-Turn Tool Use), STARE sustains stable RL training over thousands of steps while maintaining policy entropy within the target band. On AIME24 and AIME25, STARE outperforms DAPO and other competitive baselines by 4%-8% in average accuracy, with reflection tokens and response length growing in tandem, indicating sustained exploration-exploitation balance that further unlocks RL training potential.Code is available at https://github.com/hp-luo/STARE.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.