FreqForcing: 스펙트럼 자기 고정을 이용한 자동 회귀 장편 비디오 생성
FreqForcing: Autoregressive Long Video Generation via Spectral Self-Anchoring
자동 회귀 비디오 확산 모델은 실시간 스트리밍 비디오 생성을 가능하게 합니다. 그러나, 자기 순회 과정에서 발생하는 오류는 긴 시간 동안 누적되어 색상 변화, 운동 정체 및 궁극적인 시각적 붕괴를 야기합니다. 본 논문에서는 이러한 현상을 주파수 영역 관점에서 분석하여, 오류 축적이 저주파 대역에 나타나는 에너지 드리프트로 확인되었습니다. 또한, 주파수 영역에서의 어텐션 싱크의 효과를 조사한 결과, 이는 어느 정도 스펙트럼 에너지 드리프트를 완화하여 비디오 품질을 향상시키지만, 완전히 해결하지는 못함을 알 수 있었습니다. 위 분석을 바탕으로, 본 논문에서는 스펙트럼 자기 고정(SSA)을 통해 장편 비디오 생성 과정에서의 오류 누적 문제를 해결하는 훈련 불필요 프레임워크인 FreqForcing을 제안합니다. 제안된 SSA는 앵커 어텐션의 저주파 성분을 활용하여 장기적인 시각적 안정성을 유지하고, 동시에 로컬 어텐션의 고주파 성분을 통해 역동적인 움직임을 보존합니다. FreqForcing은 기존 Self-Forcing 방법을 사용하여 5초 길이의 클립으로 사전 학습된 모델을 확장하여 2분 길이의 비디오 생성이 가능하며, 이를 통해 24배에 달하는 추론 성능 향상을 달성했습니다. 광범위한 실험 결과, FreqForcing은 기존의 훈련 불필요 방법보다 양적 및 질적으로 우수한 성능을 보이며, 대표적인 훈련 기반 접근 방식과도 경쟁력 있는 수준입니다.
Autoregressive video diffusion models enable real-time streaming video generation. However, errors introduced during self-rollout accumulate over long horizons, manifesting as color drift, motion stagnation, and eventual visual collapse. In this paper, we characterize this phenomenon from a frequency-domain perspective: error accumulation appears as a pronounced energy drift in the low-frequency bands. We further investigate the effectiveness of attention sink in the frequency domain, and find that it improves the video quality by alleviating the spectral energy drift to some extent, but cannot fully resolve it. Motivated by the above analysis, we propose FreqForcing, a training-free framework that addresses error accumulation in long-video generation via Spectral Self-Anchoring (SSA). The proposed SSA leverages the low-frequency components of anchor attention to maintain long-horizon visual stability, while preserving dynamic motion through the high-frequency components of local attention. Our FreqForcing extends Self-Forcing pretrained on 5s clips to two-minute generation, achieving 24x extrapolation. Extensive experiments show that FreqForcing outperforms existing training-free methods quantitatively and qualitatively while remaining competitive with representative training-based approaches.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.