ContextMaster: 고정 예산 기반 희소 컨텍스트 라우팅을 통한 대화형 다중 샷 비디오 생성
ContextMaster: Interactive Multi-Shot Video Creation via Fixed-Budget Sparse Context Routing
최근의 비디오 모델들은 점점 더 다양한 기능을 제공하며, 단일 모델 내에서 생성, 참조 조건부 처리 및 편집 작업을 지원하지만, 일반적으로 이러한 기능은 고정된 입력에 대한 별개의 연산으로 제공됩니다. 실제 비디오 제작은 여러 샷을 포함하며, 이때 하나의 모델이 텍스트 기반 생성, 참조 추적 또는 원본 영상 편집을 수행하면서 공유되는 기록 정보를 유지해야 합니다. 본 연구에서는 이러한 설정을 대화형 다중 샷 비디오 생성(Interactive Multi-Shot Video Creation, IMVC)으로 정의하고, 각 역할에 적합한 컨텍스트 표현을 사용하는 통합 모델인 ContextMaster를 제안합니다. 대화형 모델은 확장되는 기록 정보에 대한 접근성을 유지해야 하지만, 각 디노이징 단계에서의 컨텍스트 읽기 비용 증가를 방지해야 합니다. ContextMaster는 재사용 가능한 깨끗한 컨텍스트 상태와 고정 예산 기반 희소 컨텍스트 라우팅을 결합하고, ConstraintSink를 사용하여 작업 제약 조건을 명시적으로 유지합니다. 희소 컨텍스트 접근 및 적은 디노이징 단계에서의 추론이라는 두 가지 어려움을 해결하기 위해, 우리는 두 단계로 구성된 프라이버릿 컨텍스트 증류 프레임워크를 제안합니다. 이 프레임워크는 밀집된 교사 모델로부터의 전체 컨텍스트 동작을 일관성 증류를 통해 이전하고, 배포 과정에서 분포 매칭을 사용하여 결과를 개선합니다. 세 가지 기본 작업에 대한 실험 결과는 ContextMaster가 전문화된 기준 모델보다 우수한 성능을 보이며, 다양한 작업 흐름을 유연하게 구성할 수 있음을 보여줍니다. 또한, 본 모델은 단일 GPU 환경에서 16 FPS의 속도를 달성했습니다.
Recent video models increasingly support generation, reference conditioning, and editing within a single model, yet typically expose them as separate operations over fixed inputs. Practical creation unfolds across multiple shots, requiring one model to generate from text, follow a reference, or edit source footage while maintaining shared history. We formalize this setting as interactive multi-shot video creation (IMVC) and introduce ContextMaster, a unified model with a role-aware context representation for these operations. An interactive model must retain access to an expanding history without allowing the context read cost at each denoising step to grow. ContextMaster combines reusable clean context states with fixed budget sparse context routing and uses ConstraintSink to keep task constraints visible. To address the dual challenges of sparse context access and inference with few denoising steps, we propose a two-stage privileged context distillation framework, which transfers full context behavior from a dense teacher through consistency distillation and then refines deployment rollouts with distribution matching. Experiments on the three primitive tasks demonstrate improved task fulfillment and consistency across shots over specialized baselines. User studies further validate flexibly composed workflows, while the model reaches 16 FPS on a single GPU.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.