Mage-VL: 효율적인 코덱 기반 스트리밍 멀티모달 기초 모델
Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model
기존의 대부분의 비전-언어 모델(VLM)은 모라베크 역설에 직면합니다. 즉, 복잡한 오프라인 시각적 추론에서는 뛰어난 성능을 보이지만 단순한 스트리밍 인지 작업에는 어려움을 겪으며, 처리 효율성도 떨어집니다. 본 논문에서는 실시간 멀티모달 이해 및 상호 작용을 위한 효율적인 코덱 기반 스트리밍 기초 모델인 Mage-VL을 제안합니다. 핵심 구성 요소인 Mage-ViT는 기존의 균일한 프레임 샘플링 방식을 개선하여, 모션 벡터와 희소 앵커(I) 및 예측(P) 프레임을 사용하여 동적이고 엔트로피가 높은 영역을 선택적으로 인코딩합니다. 이 모델은 16x16 패치 수준에서 작동하며, 시공간적인 맥락을 유지하면서 시각 토큰 사용량을 75% 이상 줄입니다. 약 5억 6천만 개의 비표시 이미지와 1억 개의 비디오 프레임으로 처음부터 학습된 Mage-ViT는 수십억 개의 이미지-텍스트 쌍으로 학습된 최첨단 인코더에 필적하거나 능가하는 성능을 보입니다. 또한, 멀티모달 캡셔닝을 위한 프롬프트-코드 공동 최적화 및 AI 기반 성능 진단을 포함하는 AI4AI 데이터 파이프라인을 구축했습니다. 더 나아가, 생체 영감을 받은 이중 시스템 아키텍처(가벼운 System 1 이벤트 게이트와 인과 관계를 갖는 System 2 디코더)를 통해 Mage-VL은 능동적인 스트리밍 인지를 가능하게 합니다. 광범위한 실험 결과, Mage-VL-4B는 정적 작업에서 Qwen3-VL-4B와 동등한 성능을 보이며, 비디오 이해 및 2D/3D 공간 추론 분야에서 상당한 향상을 보여줍니다. 또한, 최대 3.5배의 추론 속도 향상을 달성하고, 150억 개의 파라미터를 가진 Phi-4-reasoning-vision 기반 모델을 능가합니다. 본 연구에서는 모델 성능 외에도 사전 학습 데이터 효율성, 다양한 해상도에서의 확장 가능성, 코덱 시스템 가속화, VideoQA SFT의 중복성, 모션-공간적 상호 작용, AI4AI 데이터 파이프라인 및 멀티모달 강화 학습을 위한 Zero-Vision SFT 등 7가지 주요 경험적 결과를 제시합니다.
Standard vision-language models (VLMs) suffer from Moravec's paradox: they excel at complex offline visual reasoning but struggle with simple streaming perception tasks and process them inefficiently. We present Mage-VL, an efficient codec-native streaming foundation model for real-time multimodal understanding and interaction. At its core, our custom tokenizer, Mage-ViT, replaces uniform frame sampling by selectively encoding dynamic, entropy-rich regions using motion vectors and residual energy across sparse anchor (I) and predicted (P) frames. Operating at a 16 x 16 patch level, this reduces visual token consumption by over 75% while preserving spatiotemporal context. Trained from scratch on approximately 560M unlabeled images and 100M unlabeled video frames, Mage-ViT matches or outperforms flagship encoders trained on billions of image-text pairs. We establish AI4AI data pipelines encompassing prompt-code joint optimization for multimodal captioning and AI-driven performance diagnosis to guide training recipes. Furthermore, through a bio-inspired dual-system architecture - a lightweight System 1 event gate and a causal System 2 decoder - Mage-VL enables proactive streaming perception. Extensive evaluations show that Mage-VL-4B matches Qwen3-VL-4B on static tasks while achieving strong gains in video understanding and 2D/3D spatial reasoning, with up to a 3.5x wall-clock inference speedup, and comprehensively surpasses the 15B Phi-4-reasoning-vision baseline. Beyond model artifacts, we deliver seven key empirical findings covering pre-training data efficiency, variable-resolution scaling, codec system acceleration, VideoQA SFT redundancy, motion-spatial synergy, AI4AI data pipelines, and Zero-Vision SFT for multimodal RL.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.