2606.26917v1 Jun 25, 2026 cs.LG

GEOALIGN: 강력한 LLM 강화 학습을 위한 기하학적 Rollout 정제

GEOALIGN: Geometric Rollout Curation for Robust LLM Reinforcement Learning

Zhenqing Ling
Zhenqing Ling
Citations: 68
h-index: 4
Daoyuan Chen
Daoyuan Chen
Citations: 1,734
h-index: 20
Ting Zhou
Ting Zhou
Citations: 60
h-index: 4
Yiyang Zhao
Yiyang Zhao
Citations: 0
h-index: 0
Ying Shen
Ying Shen
Citations: 164
h-index: 7

온라인 강화 학습은 대규모 언어 모델(LLM)을 보상 신호에 맞추는 데 광범위하게 사용되지만, 노이즈가 있거나 잘못 지정된 보상은 훈련의 불안정성을 초래할 수 있습니다. 우리는 '방향 불일치'라는 실패 모드를 식별했는데, 이는 배치 내에서 소수의 높은 보상을 받는 Rollout들이 표현 공간에서의 선호 방향을 유도하며, 이 방향은 배치 다수와 크게 상충되어 높은 분산과 불안정한 업데이트를 야기합니다. 저희는 반복적 정책 최적화에서 Rollout 정제를 위한 경량 플러그인인 Geoalign을 제안합니다. Geoalign은 (i) 프롬프트 내에서의 선호 쌍을 형성하고, (ii) 각 Rollout의 숨겨진 상태에 대한 온라인 투영기를 학습하여 보상 순서대로 정렬된 변위 방향을 집중시키고, (iii) 배치 합의 프로토타입으로부터의 각도 편차를 통해 방향 불일치 Rollout을 감지하고, 프롬프트 내에서 안정적인 대체 Rollout으로 수정합니다. Geoalign은 순방향 패스만 사용하며, 무시할 수 있는 오버헤드를 추가합니다. 학습된 보상 모델과의 대화 정렬 및 이진 검증 보상을 사용하는 수학적 추론 작업에서, Geoalign은 최종 성능을 향상시키고 훈련의 진동을 줄이며, PF-PPO, PAR, PODS 및 Seed-GRPO보다 우수한 성능을 보입니다. 이러한 결과는 온라인 LLM 강화 학습에 있어 잠재적인 방향 합의가 효과적인 신뢰성 신호임을 시사합니다.

Original Abstract

Online reinforcement learning is widely used to align large language models (LLMs) with reward signals, yet training can be unstable under noisy or misspecified rewards. We identify a failure mode we call directional inconsistency: within a batch, a small set of high-reward rollouts induces representation-space preference directions that sharply disagree with the batch majority, resulting in high-variance and destabilizing updates. We propose geoalign, a lightweight plug-in for rollout curation in iterative policy optimization. Geoalign (i) forms within-prompt preference pairs, (ii) learns an online projector on per-rollout hidden states to concentrate reward-ordered displacement directions, and (iii) detects directionally inconsistent rollouts via their angular deviation from a batch consensus prototype and rectifies them with within-prompt stable alternatives. Geoalign is forward-pass only and adds negligible overhead. Across dialogue alignment with a learned reward model and mathematical reasoning with binary verified rewards, Geoalign improves final performance and reduces training oscillation, outperforming PF-PPO, PAR, PODS, and Seed-GRPO. These results suggest latent directional consensus as an effective reliability signal for online LLM RL.

0 Citations
0 Influential
10 Altmetric
50.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!