Diff-VF: 확산 모델을 이용한 학습 불필요의 고품질 장편 비디오 생성
Diff-VF: Training-free High-quality Long Video Generation via Diffusion Model
최근, 확산 모델은 비디오 생성 분야에서 상당한 발전을 이루었습니다. 그러나 대부분의 기존 비디오 확산 모델은 짧은 비디오로 훈련되었으며, 장편 비디오로 확장될 때 성능이 저하되는 경향이 있습니다. 이러한 모델들은 긴 시간 간격에서의 일관성을 유지하면서도 다양한 움직임을 표현하는 데 어려움을 겪습니다. 본 연구에서는 일관성 있고 고품질의 동적 장편 비디오를 생성하기 위해, 기존의 짧은 비디오 확산 모델을 수정하거나 미세 조정하지 않고도 이를 장편 비디오 생성기로 변환할 수 있는 학습 불필요, 플러그 앤 플레이 방식이며 모델에 독립적인 프레임워크인 Diff-VF를 제안합니다. Diff-VF는 세 가지 상호 보완적인 전략을 결합합니다. 첫째, 전역 의미를 제한하기 위한 하이브리드 노이즈 초기화(HNI), 둘째, 프레임 간의 불연속성을 제거하기 위한 가중 윈도우 샘플링(WWS), 그리고 셋째, 시간 단계에 따라 변동되는 융합을 통해 장거리 종속성을 구축하기 위한 시간 확장 샘플링(TES)입니다. 또한, Diff-VF를 Skip Residual Guidance 기술과 결합하여, 시간 단계에 따른 지침을 통해 충실도와 현실감의 균형을 맞춘 장편 비디오 품질 향상에도 활용할 수 있습니다. VBench-Long 평가 결과, Diff-VF는 FreeNoise, FreeLong 및 RIFLEx를 포함한 기존의 학습 불필요 장편 비디오 생성 모델보다 시간적 일관성과 움직임 다양성 간의 균형이 우수하며, 프레임 단위 품질 또한 경쟁력을 유지합니다. 두 가지 기본 모델에 대한 실험을 통해 다양한 공간-시간 모델링 전략을 사용하는 비디오 확산 모델에도 Diff-VF가 적용 가능하다는 것을 입증했습니다. 광범위한 추가 분석을 통해 각 구성 요소 및 하이퍼파라미터의 기여도를 검증했습니다.
Recently, diffusion models have made great progress in video generation. However, most existing video diffusion models are trained with short videos, and degrade when extrapolated to long videos, struggling to maintain long-range temporal coherence while retaining diverse motions. To generate consistent, high-quality and dynamic long videos, we propose Diff-VF, a training-free, plug-and-play and model-agnostic framework that converts existing short-video diffusion backbones into long-video generators without modifying or fine-tuning the base model. Diff-VF couples three complementary strategies: Hybrid Noise Initialization (HNI) to constrain global semantics, Weighted Window Sampling (WWS) to remove inter-window discontinuities, and Temporal Extended Sampling (TES) to establish long-range dependencies with a timestep-varying fusion. We further extend Diff-VF to long-video enhancement via Skip Residual Guidance that balances fidelity and realism through timestep-dependent guidance. VBench-Long evaluation results show that Diff-VF achieves a more favorable balance between temporal coherence and motion diversity than base models and recent training-free long video generation baselines, including FreeNoise, FreeLong, and RIFLEx, while maintaining competitive frame-wise quality. Experiments on two base models demonstrate the applicability to video diffusion models with different spatial-temporal modeling strategies. Extensive ablations validate the contribution of each component and hyperparameters.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.