2607.28058v1 Jul 30, 2026 cs.CV

롤아웃 오류로부터의 시간적 집중: 텍스트-비디오 확산 모델을 위한 암묵적 선호도 최적화

Temporal Concentration from Rollout Errors: Implicit Preference Optimization for Text-to-Video Diffusion

Kun Gai
Kun Gai
Citations: 1,042
h-index: 17
Jing Wang
Jing Wang
Citations: 50
h-index: 3
Fangyuan Kong
Fangyuan Kong
Citations: 657
h-index: 6
Xintao Wang
Xintao Wang
Citations: 2,117
h-index: 21
Henglin Liu
Henglin Liu
Citations: 21
h-index: 3
Yizhou Lin
Yizhou Lin
Citations: 6
h-index: 2
Nisha Huang
Nisha Huang
Citations: 1,105
h-index: 12
Chang Liu
Chang Liu
Citations: 19
h-index: 2
Pengfei Wan
Pengfei Wan
Citations: 7
h-index: 2
Xiu Li
Xiu Li
Citations: 75
h-index: 4

최근 확산 기반 비디오 생성, 특히 직접 선호도 최적화(DPO)를 통한 선호도 정렬 기술 발전은 시각 품질을 크게 향상시켰습니다. 그러나 시간적으로 불균일한 문제점들, 예를 들어 움직임 부재, 객체 깜빡임, 색상 과포화 현상은 여전히 현실적인 비디오 생성을 가로막는 주요 장애물입니다. 기존 방법들은 다음과 같은 두 가지 주요 한계 때문에 이러한 문제점을 해결하는 데 어려움을 겪습니다: (1) 선호도 귀속 병목 현상으로, 오프라인 인간 주석은 비용이 많이 들고 학습 동역학을 정확하게 반영하지 못하며, 온라인 보상 신호는 롤아웃 정보를 포함하지만 종종 불안정하고 편향될 수 있습니다. (2) 시간적 신용 할당 오류로, 균일하게 적용되는 감독은 문제 발생 시 짧은 구간에 효과적으로 집중하기 어렵습니다. 이러한 과제를 해결하기 위해, 우리는 비디오 확산 모델을 위한 후처리 프레임워크인 집중형 암묵적 선호도 최적화(cIPO)를 제안합니다. cIPO는 노이즈 제거 과정에서 직접 얻은 암묵적 선호도 신호를 활용합니다. 즉, 실제 비디오가 주어지면, 모델은 이를 앞으로 노이즈를 추가하고 반복적인 노이즈 제거 과정을 통해 재구성하며, 원래 비디오를 선호하는 샘플로, 재구성된 비디오를 선호하지 않는 샘플로 간주합니다. 이러한 방식은 인간 주석이나 외부 보상 모델 없이 추론 시 발생하는 오류를 포착합니다. 또한, 원본과 재구성된 비디오 간의 프레임 단위 차이는 문제가 발생한 시점을 나타냅니다. cIPO는 이러한 정보를 활용하여 시간적 재구성 오류를 계산하고 오류가 높은 구간에 최적화를 집중함으로써, 문제 발생 가능성이 높은 영역을 보다 정확하게 수정합니다. 광범위한 실험 결과에서 cIPO는 다양한 데이터셋에서 비디오의 진실성과 시간적 일관성을 지속적으로 향상시키는 것으로 나타났으며, 이는 시간적 집중을 통한 암묵적 선호도 최적화의 효과와 효율성을 강조합니다.

Original Abstract

Recent advances in preference alignment for diffusion-based video generation, particularly via Direct Preference Optimization (DPO), have significantly improved visual quality. However, temporally sparse artifacts such as motion collapse, object flickering, and color oversaturation remain a major barrier to perceptual realism. Existing methods struggle with these issues due to two key limitations: (1) the preference attribution bottleneck, where offline human annotations are costly and fail to accurately capture learning dynamics, while online reward signals are rollout-aware but often unstable and biased; and (2) temporal credit misallocation, where uniformly applied supervision cannot effectively target the brief segments in which artifacts occur. To address these challenges, we propose concentrated Implicit Preference Optimization (cIPO), a post-training framework for video diffusion models. cIPO derives implicit preference signals directly from the denoising process: given a real video, the model adds forward noise and reconstructs it via iterative denoising, treating the original as the preferred sample and the reconstruction as the dispreferred one. This formulation captures inference-time errors without requiring human annotations or external reward models. Moreover, frame-level discrepancies between original and reconstructed videos reveal when failures occur. cIPO leverages this by computing temporal reconstruction errors and concentrating optimization on high-error segments, enabling more precise correction of failure-prone regions. Extensive experiments demonstrate that cIPO consistently enhances video authenticity and temporal coherence across multiple datasets, highlighting the effectiveness and efficiency of implicit preference with temporally concentrated optimization.

0 Citations
0 Influential
11 Altmetric
55.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!