FlexComposer: 이미지부터 동영상까지 통합된 비디오 합성 프레임워크 - 유연한 경로 제어를 통한 기능
FlexComposer: Unified Video Compositing from Images to Dynamic Footage with Flexible Trajectory Control
생성적 비디오 합성은 외부 자산을 기존 비디오 시퀀스에 자연스럽게 삽입하여 콘텐츠 제작 및 시각 효과를 구현하는 데 필수적인 기술입니다. 그러나 기존 방법들은 제어 정확도와 관련된 어려움을 가지고 있습니다. 즉, 정지 이미지에서 움직임을 생성할 때 사전 애니메이션된 자산의 역동성을 유지하지 못하거나, 사용자가 정의한 경로에 따라 자산을 정확하게 배치하기 위한 세밀한 공간 제어가 부족합니다. 본 논문에서는 FlexComposer라는 통합 프레임워크를 제안합니다. FlexComposer는 비디오 합성을 경로 기반의 조건부 생성 문제로 표준화하여 정지 이미지와 동영상 모두를 원활하게 통합할 수 있도록 합니다. 저희 접근 방식은 세 가지 핵심 설계를 특징으로 합니다: (1) 객체의 고유한 움직임을 전체적인 변위로부터 분리하여 다양한 입력을 안정적이고 중앙 집중된 잠재 공간으로 표준화하는 Unified Canonical Foreground Representation; (2) VAE 잠재 공간의 병진 불변성을 활용하여 매개변수 없이 대상 경로에 핵심 특징을 전달하는 Spatial-Aware Latent Injection 전략; 그리고 (3) 절차적 시뮬레이션, 실제 영화 촬영 영상 및 생성 데이터를 결합한 Hybrid Dataset and Synthetic-to-Real Curriculum을 통해 물리적으로 타당한 조명 및 그림자 조화를 학습합니다. 이 통합 설계는 제품 사진부터 동적인 피사체까지 다양한 입력에 적용 가능하며, 명시적인 3D 재구성이나 추가적인 학습 모듈 없이도 높은 수준의 움직임 제어와 환경 통합을 달성합니다. 광범위한 실험 결과는 FlexComposer가 시각적 품질, 시간 일관성 및 경로 준수 측면에서 최첨단 방법보다 우수한 성능을 보임을 입증합니다.
Generative video compositing, which involves inserting external assets seamlessly into existing video sequences, is essential for content creation and visual effects. However, existing approaches suffer from a control-fidelity trade-off: they either hallucinate motion from static images, failing to preserve the dynamics of pre-animated assets, or lack fine-grained spatial control for precise asset placement along user-defined trajectories. We propose FlexComposer, a unified framework that standardizes video compositing as a trajectory-guided conditional generation task, enabling the seamless integration of both static images and dynamic footage. Our approach introduces three key designs: (1) a Unified Canonical Foreground Representation that decouples an object's intrinsic motion from its global displacement, standardizing heterogeneous inputs into a stabilized, centered latent space; (2) a Spatial-Aware Latent Injection strategy that exploits the translation equivariance of VAE latent spaces to transport canonical features onto target trajectories via a parameter-free mechanism; and (3) a Hybrid Dataset and Synthetic-to-Real Curriculum that synergizes procedural simulation, real-world cinematic footage, and generative data to implicitly learn physically plausible illumination and shadow harmonization. This unified design handles diverse inputs from product photos to dynamic subjects achieving high-fidelity motion control and environmental integration without the need for explicit 3D reconstruction or auxiliary learnable adapters. Extensive experiments demonstrate that FlexComposer outperforms state-of-the-art methods in visual quality, temporal consistency, and trajectory adherence.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.