컨텍스트 강제(In-Context Forcing): 자기 회귀 비디오 확산 모델에서 컨텍스트 효과 분석
In-Context Forcing: Uncovering Context Effects in Autoregressive Video Diffusion
현재의 몇 단계 자기 회귀 비디오 확산 모델은 현재 프레임의 모든 디노이징 과정에 대해 이전 프레임 전체가 완전히 노이즈 제거된 깨끗한 이미지를 컨텍스트로 사용합니다. 그러나 이러한 깨끗한 이미지에서 과도한 로컬 정보가 누출되어 모델이 단축 경로를 선택하게 만들고, 이는 시계열 의미론과 동역학을 저하시킵니다. 확산을 마스킹(masking)으로 보는 관점에서 영감을 받아, 우리는 몇 단계 자기 회귀 생성 과정에서 노이즈가 포함된 컨텍스트의 영향을 탐구합니다. 그러나 동일한 노이즈 레벨을 가진 컨텍스트를 단순히 사용하는 것만으로는 충분한 지침을 제공하지 못하며, 이는 시계열 일관성을 저해합니다. 이러한 문제를 해결하기 위해, 우리는 점진적인 자기 회귀 방식을 사용하는 '컨텍스트 강제(In-Context Forcing)'라는 새로운 방법을 제안합니다. 이 방법은 먼 프레임에는 덜, 인접한 프레임에는 더 많은 마스킹을 적용하여 적응적인 지침을 제공함으로써, 강력한 시계열 일관성과 높은 프레임 간 동역학을 동시에 확보합니다. 또한, 이전 깨끗한 이미지에 대한 엄격한 의존성을 제거함으로써, 우리 방법은 프레임 간 병렬 디노이징을 가능하게 하여 성능 저하 없이 상당한 추론 속도 향상을 달성합니다. VBench 데이터셋에 대한 광범위한 실험 결과, 제안하는 방법은 시각적 충실도와 추론 속도 측면에서 최첨단 기술보다 뛰어난 성능을 보였습니다.
Current few-step autoregressive video diffusion models depend on previous fully denoised clean frames as context for all denoising steps of the current frame. However, these clean frames leak excessive local details, which causes the model to take shortcuts, resulting in compromised temporal semantics and dynamics. Inspired by the perspective of diffusion as masking, we explore the impact of noisy contexts on few-step autoregressive generation. Yet, simply applying contexts with the same noise levels provides insufficient guidance, leading to poor temporal consistency. To resolve this dilemma, we introduce In-Context Forcing, a progressive autoregressive paradigm that utilizes contexts with decreasing noise levels. By applying less masking to distant frames and more masking to adjacent ones, this approach provides adaptive guidance, effectively ensuring both robust temporal consistency and high inter-frame dynamics. Furthermore, by decoupling the strict dependence on previous clean frames, our paradigm enables cross-frame parallel denoising, achieving substantial inference acceleration without sacrificing performance. Extensive experiments on VBench demonstrate that our method significantly outperforms state-of-the-art approaches in both visual fidelity and inference speed.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.