2607.24904v1 Jul 27, 2026 cs.CV

Mage-VL: 효율적인 코덱 기반 스트리밍 멀티모달 기초 모델

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model

Xiaoyi Zhang
Xiaoyi Zhang
Citations: 136
h-index: 6
Zhening Liu
Zhening Liu
Citations: 443
h-index: 10
Zongyu Guo
Zongyu Guo
Citations: 68
h-index: 4
Jiahao Li
Jiahao Li
Citations: 2,495
h-index: 16
Bin Li
Bin Li
Citations: 1,029
h-index: 13
Xiang An
Xiang An
Citations: 1,105
h-index: 14
Jinglu Wang
Jinglu Wang
Citations: 49
h-index: 3
Yifei Shen
Yifei Shen
Citations: 27
h-index: 3
Senqiao Yang
Senqiao Yang
Citations: 722
h-index: 9
Peng Zhang
Peng Zhang
Citations: 22
h-index: 3
Xinjie Zhang
Xinjie Zhang
Citations: 283
h-index: 9
Shicheng Zheng
Shicheng Zheng
Citations: 67
h-index: 3
J. Guo
J. Guo
Citations: 0
h-index: 0
Zhaoyang Jia
Zhaoyang Jia
Citations: 328
h-index: 8
Xun Guo
Xun Guo
Citations: 92
h-index: 3
Yuxuan Luo
Yuxuan Luo
Peking University
Citations: 14
h-index: 2
Wenxuan Xie
Wenxuan Xie
Citations: 140
h-index: 5
Kaichen Zhang
Kaichen Zhang
Citations: 98
h-index: 3
Zihan Zheng
Zihan Zheng
Citations: 19
h-index: 3
Xiao Li
Xiao Li
Citations: 367
h-index: 11
Yan Lu
Yan Lu
Citations: 2,157
h-index: 16
Haoqing Wang
Haoqing Wang
Citations: 582
h-index: 8
Yin Xie
Yin Xie
Citations: 216
h-index: 4

기존의 대부분의 비전-언어 모델(VLM)은 모라베크 역설에 직면합니다. 즉, 복잡한 오프라인 시각적 추론에서는 뛰어난 성능을 보이지만 단순한 스트리밍 인지 작업에는 어려움을 겪으며, 처리 효율성도 떨어집니다. 본 논문에서는 실시간 멀티모달 이해 및 상호 작용을 위한 효율적인 코덱 기반 스트리밍 기초 모델인 Mage-VL을 제안합니다. 핵심 구성 요소인 Mage-ViT는 기존의 균일한 프레임 샘플링 방식을 개선하여, 모션 벡터와 희소 앵커(I) 및 예측(P) 프레임을 사용하여 동적이고 엔트로피가 높은 영역을 선택적으로 인코딩합니다. 이 모델은 16x16 패치 수준에서 작동하며, 시공간적인 맥락을 유지하면서 시각 토큰 사용량을 75% 이상 줄입니다. 약 5억 6천만 개의 비표시 이미지와 1억 개의 비디오 프레임으로 처음부터 학습된 Mage-ViT는 수십억 개의 이미지-텍스트 쌍으로 학습된 최첨단 인코더에 필적하거나 능가하는 성능을 보입니다. 또한, 멀티모달 캡셔닝을 위한 프롬프트-코드 공동 최적화 및 AI 기반 성능 진단을 포함하는 AI4AI 데이터 파이프라인을 구축했습니다. 더 나아가, 생체 영감을 받은 이중 시스템 아키텍처(가벼운 System 1 이벤트 게이트와 인과 관계를 갖는 System 2 디코더)를 통해 Mage-VL은 능동적인 스트리밍 인지를 가능하게 합니다. 광범위한 실험 결과, Mage-VL-4B는 정적 작업에서 Qwen3-VL-4B와 동등한 성능을 보이며, 비디오 이해 및 2D/3D 공간 추론 분야에서 상당한 향상을 보여줍니다. 또한, 최대 3.5배의 추론 속도 향상을 달성하고, 150억 개의 파라미터를 가진 Phi-4-reasoning-vision 기반 모델을 능가합니다. 본 연구에서는 모델 성능 외에도 사전 학습 데이터 효율성, 다양한 해상도에서의 확장 가능성, 코덱 시스템 가속화, VideoQA SFT의 중복성, 모션-공간적 상호 작용, AI4AI 데이터 파이프라인 및 멀티모달 강화 학습을 위한 Zero-Vision SFT 등 7가지 주요 경험적 결과를 제시합니다.

Original Abstract

Standard vision-language models (VLMs) suffer from Moravec's paradox: they excel at complex offline visual reasoning but struggle with simple streaming perception tasks and process them inefficiently. We present Mage-VL, an efficient codec-native streaming foundation model for real-time multimodal understanding and interaction. At its core, our custom tokenizer, Mage-ViT, replaces uniform frame sampling by selectively encoding dynamic, entropy-rich regions using motion vectors and residual energy across sparse anchor (I) and predicted (P) frames. Operating at a 16 x 16 patch level, this reduces visual token consumption by over 75% while preserving spatiotemporal context. Trained from scratch on approximately 560M unlabeled images and 100M unlabeled video frames, Mage-ViT matches or outperforms flagship encoders trained on billions of image-text pairs. We establish AI4AI data pipelines encompassing prompt-code joint optimization for multimodal captioning and AI-driven performance diagnosis to guide training recipes. Furthermore, through a bio-inspired dual-system architecture - a lightweight System 1 event gate and a causal System 2 decoder - Mage-VL enables proactive streaming perception. Extensive evaluations show that Mage-VL-4B matches Qwen3-VL-4B on static tasks while achieving strong gains in video understanding and 2D/3D spatial reasoning, with up to a 3.5x wall-clock inference speedup, and comprehensively surpasses the 15B Phi-4-reasoning-vision baseline. Beyond model artifacts, we deliver seven key empirical findings covering pre-training data efficiency, variable-resolution scaling, codec system acceleration, VideoQA SFT redundancy, motion-spatial synergy, AI4AI data pipelines, and Zero-Vision SFT for multimodal RL.

0 Citations
0 Influential
8 Altmetric
40.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!