2607.29545v1 Jul 31, 2026 cs.CV

MoRoute: 문맥 내 다중 모드 비디오 생성 시스템을 위한 동적 라우팅

MoRoute: Dynamic Routing for In-Context Multimodal Video Generation

Zhan Peng
Zhan Peng
Citations: 75
h-index: 5
Jun Liang
Jun Liang
Citations: 37
h-index: 2
Jing Li
Jing Li
Citations: 85
h-index: 5
Jie Ma
Jie Ma
Citations: 17
h-index: 1
Haoxue Wu
Haoxue Wu
Citations: 93
h-index: 4
Guanbin Li
Guanbin Li
Citations: 3,856
h-index: 32
Chong Gao
Chong Gao
Citations: 0
h-index: 0
Chongxiao Wang
Chongxiao Wang
Citations: 0
h-index: 0

다중 모드 비디오 생성은 단일 모델 내에서 텍스트, 이미지 및 비디오의 다양한 조합에 따라 비디오를 생성하고 편집하는 것을 목표로 하며, 이를 통해 다양한 작업이 상호 보완적인 데이터와 생성 규칙을 공유할 수 있습니다. 이러한 작업을 통합하려면 다양한 조건에 대한 다중 모드 이해가 필요하며, 이는 일반적으로 사전 훈련된 시각-언어 모델(VLM)을 통해 제공됩니다. 주요 과제는 VLM의 계층적 다중 모드 표현과 사전 훈련된 비디오 디퓨전 트랜스포머(DiT)를 어떻게 연결할 것인가입니다. 기존 방법은 VLM의 최종 레이어나 몇 개의 수동으로 선택한 레이어에서만 특징을 주입하거나, 구조적으로 일치하는 이해 및 생성 스트림을 공동으로 학습하여, 이질적인 사전 훈련된 백본을 재사용하기 어렵게 만듭니다. 본 논문에서는 동적 레이어 라우팅을 통해 서로 다른 아키텍처를 가진 고정된 VLM과 사전 훈련된 비디오 DiT를 연결하는 통합 다중 모드 비디오 생성 프레임워크인 MoRoute를 소개합니다. 각 입력에 대해, 경량 블록 단위 라우터가 각 DiT 블록이 자신의 생성 단계에 가장 관련성이 높은 VLM 레이어를 선택하도록 하여, 다중 모드 이해와 비디오 합성을 위한 적응적 대응 관계를 학습합니다. 또한, MoRoute는 참조 이미지와 원본 비디오를 통합된 문맥 조건화를 통해 DiT 토큰 시퀀스에 직접 통합하여, 다양한 생성 및 편집 작업에서 미세한 시각적 세부 사항을 유지합니다. IntelligentVBench, OpenVE-Bench 및 RefVIE-Bench에서의 실험 결과, MoRoute는 각 벤치마크에서 가장 뛰어난 성능을 보였으며, 평균 점수를 각각 0.15, 0.18 및 0.34만큼 향상시켰습니다 (1~5 척도 기준).

Original Abstract

Multimodal video generation aims to generate and edit videos conditioned on arbitrary combinations of text, images, and videos within a single model, allowing diverse tasks to share complementary data and generative priors. Unifying these tasks requires multimodal understanding of diverse conditions, which is typically provided by a pretrained vision-language model (VLM). A key challenge is how to connect the VLM's hierarchical multimodal representations with a pretrained video diffusion transformer (DiT). Existing methods either inject features from only the final or a few manually selected VLM layers, or jointly train architecture-matched understanding and generation streams, making it difficult to reuse heterogeneous pretrained backbones. We introduce MoRoute, a unified multimodal video generation framework that formulates a frozen VLM and a pretrained video DiT with different architectures as heterogeneous experts connected through dynamic layer routing. For each input, a lightweight block-wise router enables every DiT block to select the VLM layer most relevant to its generation stage, thereby learning an adaptive correspondence between multimodal understanding and video synthesis. MoRoute further incorporates reference images and source videos directly into the DiT token sequence through unified in-context conditioning, preserving fine-grained visual details across diverse generation and editing tasks. Experiments on IntelligentVBench, OpenVE-Bench, and RefVIE-Bench show that MoRoute consistently surpasses the best competing method on each benchmark, improving the average score by 0.15, 0.18, and 0.34 on a 1-5 scale, respectively.

0 Citations
0 Influential
16 Altmetric
80.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!