CineWeaver: 학습 없이 참조 기반의 다중 프레임 장편 비디오 생성을 통한 영화적 스토리텔링
CineWeaver: Training-Free Reference-Controllable Multi-Shot Long Video Generation for Cinematic Storytelling
텍스트-비디오 확산 모델은 여러 장면을 동시에 생성하고, 캐릭터 및 장면의 미세한 제어를 제공하며, 긴 시간 동안 지속되는 영상을 생성해야 하므로 영화 제작용 비디오 생성이 어렵습니다. 기존 방법들은 이러한 특정 요구 사항을 개별적으로 해결하기 위해 커스터마이징과 재학습에 의존하며, 하나의 통합 프레임워크로 모든 요구 사항을 동시에 충족할 수 없습니다. 본 논문에서는 학습이 필요 없는 패러다임을 제시하고, 다중 프레임 생성이 어려운 이유가 사전 훈련된 비디오 확산 모델의 시간적 연속성에 대한 구조적 편향 때문이라는 핵심 통찰력을 바탕으로, 참조 기반의 다중 프레임 장편 비디오 생성을 위한 통합 프레임워크인 CineWeaver를 제안합니다. 우리는 추론 과정에서 위치 인코딩 및 어텐션 패턴을 조작하여 시간적 연속성을 깨뜨려 사전 훈련된 비디오 확산 모델을 사용하여 명확한 장면 전환을 가능하게 합니다. 또한, 제안된 프레임워크에 샷 단위의 미세한 제어를 위한 참조 기반 조건부 메커니즘을 확장하고, 일관된 전반적인 외형 단서를 활용하여 장편 생성을 가능하게 하는 앵커 메모리 메커니즘을 개발했습니다. 현재까지 알려진 바로는, CineWeaver는 extbf{장편}, extbf{참조 기반} 및 extbf{다중 프레임} 비디오 생성을 동시에 가능하게 하는 첫 번째 통합 프레임워크입니다. 실험 결과는 CineWeaver가 일관된 개체 인식, 안정적인 전반적인 외형 및 명확한 장면 전환을 갖춘 고품질의 장편 영화 영상을 생성함을 보여줍니다. 프로젝트 페이지는 다음에서 확인할 수 있습니다: https://cineweaver.github.io.
Cinematic video generation is challenging for text-to-video diffusion models due to concurrent requirements on multi-shot generation, fine-grained controllability over characters and scenes, and long-form generation across extended temporal horizons. Existing methods rely on customization and retraining to separately address specific requirements, and cannot simultaneously fulfill all the requirements with a unified framework. In this paper, we shed light on the training-free paradigm with the key insight that the difficulty of multi-shot generation arises from a structural bias toward temporal continuity in pretrained video diffusion models, and consequently, propose a unified framework named CineWeaver to achieve reference-controllable multi-shot long-video generation without retraining. We manipulate positional encoding and attention patterns to break temporal continuity during inference to enable clear shot transitions using pretrained video diffusion models. Furthermore, we extend the proposed framework with a shot-routed reference conditioning mechanism for per-shot fine-grained controllability, and develop an anchor memory mechanism to allow long-form generation with consistent global appearance cues. To our best knowledge, CineWeaver is the first unified framework to simultaneously enable \textbf{long-form}, \textbf{reference-controllable}, and \textbf{multi-shot} video generation in a training-free fashion. Experimental results demonstrate that CineWeaver produces high-quality cinematic videos of long durations with consistent identities, stable global appearance, and clear shot transitions. The project page is available at: https://cineweaver.github.io.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.