확산 모델의 선호도 정렬을 위한 잠재 보상 레지스터
Latent Reward Registers for Diffusion Preference Alignment
확산 모델을 인간의 선호도에 맞추는 것은 일반적으로 최종 생성 결과물에 대한 희소한 종료 보상을 활용하는데, 이는 다단계 노이즈 제거 과정 전반에 걸쳐 심각한 시간적 신용 할당 문제를 야기합니다. 본 논문에서는 Latent Reward Registers라는 메커니즘을 제안합니다. 이 메커니즘은 고정된 Diffusion Transformer (DiT)의 입력 시퀀드 앞에 학습 가능한, 위치에 구애받지 않는 레지스터 토큰을 추가하여 중간 단계의 노이즈가 있는 잠재 변수로부터 직접 종료 선호도를 추정합니다. 이러한 독립적인 읽기 메커니즘은 생성기의 숨겨진 상태나 속도장을 변경하지 않고 잠재 보상 정보를 추출합니다. 결과적으로, 전체 노이즈 제거 과정에서 밀도가 높고 미분 가능한 보상 신호를 제공하여 두 가지 정렬 전략을 가능하게 합니다. 학습 단계에서는 Reward-Gradient On-Policy Distillation (RG-OPD)이라는 방법이 정책 경사 방법의 계산 비용이 많이 드는 시뮬레이션을 거치지 않고, 보상을 활용한 업데이트를 on-policy 경로에 적용합니다. 추론 단계에서는 Reward-Guided Sampling (RGS)이라는 방법이 파라미터 업데이트 없이 보상 기반 기울기를 사용하여 경로를 조정합니다. 실험 결과, 높은 노이즈 수준(u = 0.8)에서 제안하는 레지스터는 평가된 다른 잠재 보상 모델 중에서 가장 높은 쌍별 정확도를 달성했습니다. 또한, RG-OPD는 기존 강화 학습 방법보다 우수한 성능을 보이며 GPU 사용 시간을 최대 33배까지 줄입니다. RGS는 학습 없이 수행되는 방법 중 최고 수준의 성과를 보여주며, 정렬 및 시각적 품질 모두에서 상당한 개선을 가져옵니다. 코드와 모델 가중치는 https://github.com/Guanys-dar/latent-reward-register 에서 확인할 수 있습니다.
Aligning diffusion models with human preferences usually relies on a sparse terminal reward evaluated on the final generated samples, presenting a severe temporal credit-assignment challenge across the multi-step denoising process. We propose Latent Reward Registers, a mechanism that estimates terminal preference directly from intermediate noisy latents by prepending learnable, position-free register tokens to the input sequence of a frozen Diffusion Transformer (DiT). This independent readout mechanism extracts latent reward evidence without altering the generator's hidden states or velocity field. The resulting dense, differentiable reward signal throughout the full denoising process facilitates two alignment strategies. For training, Reward-Gradient On-Policy Distillation (RG-OPD) distills reward-guided updates along on-policy trajectories, bypassing the computationally expensive rollouts of standard policy gradients. For inference, Reward-Guided Sampling (RGS) steers trajectories via magnitude-matched reward gradients without parameter updates. Empirically, at high noise levels (u = 0.8), the registers reach the highest pairwise accuracy among the evaluated latent reward models. Furthermore, RG-OPD outperforms online reinforcement learning baselines while reducing GPU hours by up to 33x, and RGS establishes a new state-of-the-art among training-free methods, strictly enhancing both alignment and perceptual metrics. Code and weights are available at https://github.com/Guanys-dar/latent-reward-register
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.