HPO: 희스테리 정책 최적화 - 희소 보상 환경에서의 안정적이고 효율적인 학습을 위한 방법
HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime
본 연구에서는 검증 가능한 희소 보상을 사용하는 GRPO 스타일 강화 학습에서 발생하는 특정 유형의 실패 현상을 분석합니다. 초기 업데이트 단계에서 양의 이점을 가진 응답보다 음의 이점을 가진 응답이 더 많게 나타나는 경향이 있으며, 응답 수준의 길이 정규화는 업데이트 크기를 출력 길이에 연결하여 문제를 심화시킵니다. 우리는 이러한 문제를 해결하기 위해 GRPO를 최소한으로 수정하고 음의 이점 업데이트 가중치를 줄이며, 응답별 길이 정규화를 평균 길이 정규화로 대체하는 희스테리 정책 최적화 (HPO) 방법을 제안합니다. 또한, 배치 수준의 이점 부호 통계를 기반으로 희스테리 가중치를 조절하는 적응형 HPO (A-HPO)를 도입하여 고정된 희스테리 가중치 튜닝의 필요성을 없앱니다. TeleLogs 및 Countdown 실험에서 A-HPO는 GRPO에 비해 업데이트당 보상을 향상시켰으며, 특히 초기 희소 보상 환경에서 더 큰 성능 향상을 보였습니다. TeleLogs에서는 A-HPO가 SAPO보다 5%, GSPO보다 11%, GRPO보다 15% 높은 최종 보상인 0.84를 달성했으며, 응답 길이는 유사한 수준을 유지했습니다. Countdown에서는 A-HPO가 1.5B에서 7B 모델 크기 범위에서 초기 및 가장 어려운 구성에서 가장 큰 성능 향상을 보여주었습니다. 희스테리 가중치에 대한 분석 결과는 A-HPO의 성능 향상이 양전적인 업데이트만 사용하거나 완전히 대칭적인 업데이트를 사용하는 방식보다 양성 및 음성의 이점 기여도를 더 균형 있게 조절하기 때문에 발생한다는 것을 나타냅니다.
We investigate a narrow but common failure mode of GRPO-style reinforcement learning in the context of sparse verifiable rewards: early updates contain more responses with negative advantages than those with positive advantages, while response-level length normalization ties the magnitude of the update to the length of the output. We propose Hysteretic Policy Optimization (HPO), a minimal modification of GRPO that reduces the weight of negative-advantage updates and replaces per-response length normalization with mean-length normalization. We further introduce Adaptive HPO (A-HPO), which sets the hysteretic weight based on batch-level advantage-sign statistics, thereby removing the need for tuning a fixed hysteretic weight. In our TeleLogs and Countdown experiments, A-HPO improves the reward per update compared to GRPO, with the largest gains in early sparse reward regimes. On TeleLogs, A-HPO achieves a final reward of 0.84, outperforming SAPO by 5%, GSPO by 11%, and GRPO by 15%, while maintaining a comparable response-length. On Countdown, A-HPO achieves the largest gains in initial and most difficult configurations across 1.5B-7B models. Ablation studies on the hysteretic weight show that the gains of A-HPO come from better balancing the contributions of positive and negative advantages compared to positive-only or fully symmetric updates.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.