AdvantageFlow: 강화 학습에서 플로우 모델을 위한 가중 최소 제곱법
AdvantageFlow: Advantage-Weighted Least Squares for RL in Flow Models
본 논문에서는 수정된 플로우 모델을 위한 순방향 프로세스 강화 학습 알고리즘인 AdvantageFlow를 소개합니다. Flow-GRPO와 달리, 역방향 프로세스를 최적화하는 대신, 가중치를 적용한 순방향 프로세스 예측 손실을 최적화합니다. 이 최적화 문제는 긍정적인 어드밴티지가 없을 때 불안정해지고, 손실 함수가 비볼록하게 변할 수 있습니다. 우리는 로우트 정책 정규화를 통해 이를 안정화하며, 이는 분산을 줄이고, 지역적으로 보상을 향상시키는 목표 분포에 적합하는 방식으로 발생합니다. AdvantageFlow는 Stable Diffusion 3.5 Medium을 사용한 이미지 생성 작업에서 평가되었으며, Flow-GRPO 및 음수 인식 미세 조정 기반의 최첨단 순방향 프로세스 강화 학습 기준 모델보다 우수한 성능을 보였습니다.
We introduce AdvantageFlow, a forward-process reinforcement learning algorithm for rectified flow models. Unlike Flow-GRPO, which optimizes the reverse process, we optimize an advantage-weighted forward-process prediction loss. This optimization problem is unstable when advantages are negative and the loss becomes non-convex. We stabilize it by rollout policy regularization, which reduces variance and arises from fitting a local reward-improving target distribution. We evaluate AdvantageFlow on image generation tasks with Stable Diffusion 3.5 Medium. It outperforms both Flow-GRPO and a state-of-the-art forward-process RL baseline based on negative-aware fine-tuning.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.