2608.06125v1 Aug 06, 2026 cs.CV

불확실성 기반 확산 모델 후처리 학습을 위한 샘플 적응형 잠재 보상

Sample-Adaptive Latent Rewards for Uncertainty-Guided Diffusion Post-Training

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Chi Zhang
Chi Zhang
Citations: 52
h-index: 4
Xuelong Li
Xuelong Li
Citations: 8
h-index: 2
Kechun Hao
Kechun Hao
Citations: 1
h-index: 1
Ziqiao Weng
Ziqiao Weng
Citations: 59
h-index: 4

잠재 보상 모델은 중간 상태를 픽셀 공간으로 디코딩하지 않고도 시각적 확산 모델을 지도할 수 있으며, 이는 인간의 선호도와의 일관성을 더욱 효율적으로 만듭니다. 그러나 기존의 잠재 보상 모델은 스칼라 점수만을 출력하며, 각 예측의 불확실성을 추정하지 않습니다. 따라서 생성자는 어떤 피드백이 신뢰할 수 있는지 판단할 수 없습니다. 이는 최적화가 잘못된 방향으로 진행되도록 유도하고, 보상 해킹을 초래할 수 있습니다. 본 논문에서는 이미지 및 비디오 확산 모델에 대한 통일된 잠재 공간 프레임워크인 extsc{SURE}를 제안합니다. 이는 보상 분포를 학습하고, 이러한 분포의 신뢰도를 직접 사용하여 밀집적인 후처리 학습을 안내합니다. 먼저, 샘플 적응형 잠재 보상 모델 ( extsc{SURE-LRM})을 제안합니다. 이 모델은 각 노이즈가 섞인 잠재 변수에 대해 가우시안 유틸리티를 예측하며, 그 평균값은 보상 점수를 나타내고, 분산은 인간의 주석 없이 예측의 불확실성을 반영합니다. 학습된 분포는 불확실성 기반 보상 피드백 학습 ( extsc{SURE-REFL})을 통해 후처리 학습을 안내합니다. 이 방법은 디노이징 경로를 따라 불확실성 기반의 밀집적인 피드백을 제공합니다. 선택된 특정 시점에서, extsc{SURE-REFL}은 고정된 extsc{SURE-LRM}에 쿼리를 수행하고, 분산을 계산하여 동일한 시점에서의 샘플에 대한 신뢰도 가중치를 부여합니다. 각 가중치가 적용된 보상은 해당 로컬 트랜지션만을 통해 역전파됩니다. 전체 프로세스는 잠재 공간 내에서 이루어지며, 픽셀 공간 디코딩이나 전체 디노이징 그래프가 필요하지 않습니다. 실험 결과, extsc{SURE-LRM}은 강력한 기준 모델보다 선호도 예측 성능을 향상시켰습니다. extsc{SURE-REFL}은 다양한 지표에서 최고 성능을 달성했으며, 최적화 안정성을 더욱 개선했습니다. 또한, 평가된 방법 중 가장 높은 VBench 품질, 의미론 및 총점을 기록했습니다.

Original Abstract

Latent reward models can supervise visual diffusion models without decoding intermediate states into pixel space. This makes alignment with human preferences more efficient. However, existing latent reward models output only scalar scores. They do not estimate the uncertainty of each prediction. The generator therefore cannot determine which feedback is reliable. This can drive optimization in the wrong direction and lead to reward hacking. We propose \textsc{SURE}, a unified latent-space framework for image and video diffusion models. It learns reward distributions and directly uses their reliability to guide dense post-training. First, we propose sample-adaptive latent reward model (\textsc{SURE-LRM}). It predicts a Gaussian utility for each noisy latent. Its mean predicts the reward score. Its variance reflect the uncertainty of prediction without human annotation. The learned distribution then guides post-training through uncertainty-guided reward feedback learning (\textsc{SURE-REFL}). This method provides uncertainty-guided dense feedback along the denoising trajectory. At selected transitions, \textsc{SURE-REFL} queries the frozen \textsc{SURE-LRM}. It converts detached variance into reliability weights for samples at the same transition. Each weighted reward is backpropagated only through its local transition. The entire process remains in latent space and requires neither pixel-space decoding nor the full denoising graph. Experiments show that \textsc{SURE-LRM} improves preference prediction over strong baselines. \textsc{SURE-REFL} achieves the sota performance among various metrics and further improves optimization stability. It also achieves the highest VBench quality, semantic, and total scores among the evaluated methods.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!