2601.00423v1 Jan 01, 2026 cs.LG

E-GRPO: 높은 엔트로피 단계가 흐름 모델에 대한 효과적인 강화 학습을 이끄는 방식

E-GRPO: High Entropy Steps Drive Effective Reinforcement Learning for Flow Models

Shengjun Zhang
Shengjun Zhang
Citations: 150
h-index: 6
Zhang Zhang
Zhang Zhang
Citations: 26
h-index: 3
Chensheng Dai
Chensheng Dai
Citations: 17
h-index: 2
Yueqi Duan
Yueqi Duan
Citations: 13
h-index: 1

최근 강화 학습은 흐름 매칭 모델의 인간 선호도 정렬을 향상시키는 데 기여해 왔습니다. 확률적 샘플링은 디노이징 방향을 탐색하는 데 도움을 주지만, 여러 디노이징 단계를 최적화하는 기존 방법은 희소하고 모호한 보상 신호로 인해 어려움을 겪습니다. 우리는 높은 엔트로피 단계가 더욱 효율적이고 효과적인 탐색을 가능하게 하지만, 낮은 엔트로피 단계는 구별되지 않는 결과를 초래한다는 것을 관찰했습니다. 이에 따라, 우리는 SDE 샘플링 단계의 엔트로피를 증가시키는 엔트로피 기반 그룹 상대 정책 최적화(E-GRPO)를 제안합니다. 확률 미분 방정식(SDE) 통합은 여러 단계에서 발생하는 확률성으로 인해 모호한 보상 신호를 발생시키므로, 우리는 연속된 낮은 엔트로피 단계를 하나의 높은 엔트로피 단계로 병합하여 SDE 샘플링을 수행하고, 다른 단계에서는 상미분 방정식(ODE) 샘플링을 적용합니다. 이를 바탕으로, 우리는 동일한 통합된 SDE 디노이징 단계를 공유하는 샘플 내에서 그룹 상대적인 장점을 계산하는 다단계 그룹 정규화 장점을 도입했습니다. 다양한 보상 설정에 대한 실험 결과는 제안하는 방법의 효과를 입증했습니다.

Original Abstract

Recent reinforcement learning has enhanced the flow matching models on human preference alignment. While stochastic sampling enables the exploration of denoising directions, existing methods which optimize over multiple denoising steps suffer from sparse and ambiguous reward signals. We observe that the high entropy steps enable more efficient and effective exploration while the low entropy steps result in undistinguished roll-outs. To this end, we propose E-GRPO, an entropy aware Group Relative Policy Optimization to increase the entropy of SDE sampling steps. Since the integration of stochastic differential equations suffer from ambiguous reward signals due to stochasticity from multiple steps, we specifically merge consecutive low entropy steps to formulate one high entropy step for SDE sampling, while applying ODE sampling on other steps. Building upon this, we introduce multi-step group normalized advantage, which computes group-relative advantages within samples sharing the same consolidated SDE denoising step. Experimental results on different reward settings have demonstrated the effectiveness of our methods.

13 Citations
2 Influential
3 Altmetric
32.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!