2601.03895v1 Jan 07, 2026 cs.LG

적응적 경계 클리핑 GRPO: 안정적이고 일반화 가능한 학습을 위한 경계 비율 보장

All-Quadrant Bounded Clipping GRPO: Closing the Unbounded Blind Spot for Stable and Generalizable Training

Xin Chen
Xin Chen
Citations: 93
h-index: 5
C. Liu
C. Liu
Citations: 46
h-index: 1

그룹 상대 정책 최적화(GRPO)는 대규모 언어 모델(LLM)을 활용한 강화 학습에서 인기 있는 알고리즘으로 부상했습니다. 그러나 GRPO의 클리핑 메커니즘을 분석한 결과, 특정 시나리오에서 최적이 아니라는 점을 발견했습니다. 적절한 수정 사항을 적용하면 GRPO를 크게 개선하여 유연성과 일반화 성능을 향상시킬 수 있습니다. 이에 따라, 본 논문에서는 원래 GRPO 프레임워크의 비대칭적이고 적응적인 개선 버전인 Adaptive-Boundary-Clipping GRPO (ABC-GRPO)를 제안합니다. 우리는 Qwen3 LLM을 사용한 수학적 추론 작업에서 ABC-GRPO가 표준 GRPO보다 우수한 성능을 달성한다는 것을 보여줍니다. 또한, ABC-GRPO는 훈련 과정 전반에 걸쳐 훨씬 높은 엔트로피를 유지하여 모델의 탐색 능력을 보존하고 조기 수렴을 완화합니다. 구현 코드는 온라인에서 제공되어 재현성을 돕습니다: https://github.com/chi2liu/ABC-GRPO.

Original Abstract

Group Relative Policy Optimization (GRPO) has emerged as a popular algorithm for reinforcement learning with large language models (LLMs). However, GRPO inherits PPO's token-level clipping while replacing token-level advantages with a single sequence-level advantage. Through a four-quadrant analysis of the (likelihood-ratio, advantage) space, we show that this combination leaves one quadrant -- negative advantage combined with an increased likelihood ratio (Q4) -- structurally unbounded, so that a few high-ratio tokens can receive very large suppressive updates that collapse entropy and narrow the reasoning boundary. To address this, we propose All-Quadrant Bounded Clipping GRPO (ABC-GRPO), which applies unconditional clipping in all four quadrants through sign-dependent boundaries. ABC-GRPO clips the likelihood ratio before multiplying by the advantage, adding a trust-region floor in Q2 and a cap in Q4 -- its negative-advantage branch coinciding with dual-clip PPO -- to yield bounded per-step policy displacement in every quadrant. On mathematical reasoning with Qwen3 base models, ABC-GRPO attains the highest Avg@64 and Pass@64: it is statistically superior to GRPO, SAPO, and dual-clip PPO and competitive with the strongest baseline (DAPO), while maintaining substantially higher entropy; the gains transfer to MATH-500 and to out-of-domain code (HumanEval). Ablations isolate Q4 as the dominant blind spot.

0 Citations
0 Influential
39.982537807332 Altmetric
0.0 Score
Original PDF
32

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!