DiT-Reward: 텍스트-이미지 생성 모델의 보상 모델링을 위한 생성 표현
DiT-Reward: Generative Representations for Text-to-Image Reward Modeling
이미지 생성을 위해 학습된 표현은 생성된 이미지의 평가에도 활용될 수 있을까요? 본 연구에서는 생성 표현 학습의 하위 작업으로 텍스트-이미지 보상 예측을 탐구합니다. 이를 위해, 사전 학습된 텍스트-이미지 디퓨전 트랜스포머를 'DiT-Reward'라는 보상 모델로 변환하는 방법을 제시합니다. DiT-Reward는 깨끗한 이미지의 잠재 공간 표현을 처리하고, 트랜스포머 레이어 전반에 걸쳐 텍스트 조건부 이미지 표현을 결합하여 작동합니다. 동일한 학습 데이터셋 하에서, DiT-Reward는 HPSv3보다 모든 평가 기준점(preference benchmarks)에서 더 우수한 성능을 보이며, 특히 HPDv2에서 85.6%, HPDv3에서 77.6%의 정확도를 달성했습니다. 생성 모델의 기본 구조를 고정했을 때에도, 경량화된 학습 헤드는 여전히 표현으로부터 의미 있는 선호도 예측을 추출할 수 있습니다. 깊이 방향으로 분석한 결과, 하위 작업에서의 보상 성능은 중간-후반 레이어에서 가장 강력하며, 다양한 단계의 표현을 결합하는 것이 효과적입니다. 또한, 생성 모델의 기본 구조 용량과 긍정적인 상관 관계가 있음을 확인했습니다. 마지막으로, Flow-GRPO를 사용하여 Stable Diffusion 3.5 Large 모델을 최적화할 때, DiT-Reward는 동일한 학습 경로에서 HPSv3보다 우수한 성능을 보이며, 특히 현실감 측면에서 뚜렷한 개선이 나타났습니다. 또한, 잠재 공간 직접 점수화 방식은 최대 메모리 사용량은 비슷하면서도 추론 속도를 HPSv3에 비해 1.65배 향상시켰습니다. 이러한 결과는 사전 학습된 생성 트랜스포머가 보상 모델링 및 정책 최적화를 위한 전이 가능한 표현을 제공할 수 있음을 보여줍니다.
Can representations learned for image generation also support the evaluation of generated images? We study text-to-image reward prediction as a downstream task of generative representation learning. To this end, we introduce DiT-Reward, which converts a pretrained text-to-image Diffusion Transformer into a reward model by processing near-clean image latents and aggregating text-conditioned image representations across transformer layers. Under the same training data mixture as HPSv3, DiT-Reward outperforms HPSv3 on all four evaluated preference benchmarks, reaching 85.6% on HPDv2 and 77.6% on HPDv3. When the generative backbone is frozen, a lightweight learned head can still extract meaningful preference predictions from its representations. Probing across depth further reveals that downstream reward performance is strongest in the middle-to-late layers and benefits from combining representations across different stages. We also observe consistent positive scaling with generative backbone capacity. Finally, when used to optimize Stable Diffusion 3.5 Large with Flow-GRPO, DiT-Reward outperforms HPSv3 along the matched training trajectory, with particularly clear gains in realism. Direct latent scoring also achieves a 1.65x inference speedup over HPSv3 with comparable peak memory. These results show that pretrained generative DiTs provide transferable representations for reward modeling and policy optimization.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.