다중 모드 사전 학습의 물리학적 이해를 향하여: 지식 흐름, 모달리티 시너지, 초기 통합 및 실용적인 방법론
Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes
컴퓨터 비전은 기초 모델 발전에 중요한 역할을 하며, 이를 통해 다중 모드를 통합한 사전 학습으로의 전환을 주도하고 있습니다. 하지만, 이러한 발전에도 불구하고, 통일된 학습 과정에서 다양한 모달리티 간의 상호 작용 방식과 설계 공간에 대한 근본적인 이해는 아직 부족합니다. 본 연구에서는 체계적인 탐구를 통해 다중 모드 사전 학습에 대한 경험적 증거를 제공합니다. 합성 데이터와 대규모 실제 데이터를 사용한 통제된 실험을 통해, 다중 모드 사전 학습의 '물리학'에 대한 네 가지 핵심적인 인사이트를 얻었습니다: (i) 지식 흐름: 언어, 시각적 이해 및 시각적 생성 과정에서 각 모달리티가 어떻게 서로 지식을 전달하는지 분석하여, 영향력과 비대칭성이 뚜렷한 패턴을 밝혀냈습니다; (ii) 시너지와 경쟁: 데이터의 '복잡성'이 모달리티 간의 상호 작용이 시너지를 내는지 경쟁적인 관계를 보이는지를 결정하며, 공유 어텐션 및 정규화와 같은 특정 아키텍처 선택이 시너지를 촉진한다는 것을 확인했습니다. 이러한 경향은 다양한 시각적 토크나이저 설계에서도 나타납니다; (iii) 초기 통합: 모달리티를 초기 단계부터 통합하여 동시에 학습하는 것이, 나중에 정렬하거나 순차적으로 학습하는 것보다 효과적입니다. 이 과정에서 '시각적 게으름' 현상이 발견되었는데, 이는 지연된 통합이 모델을 언어 정보에 과도하게 의존하게 만드는 것을 의미합니다; (iv) 실용적인 방법론: 전체 연산 자원의 5%만을 사용하여 강력한 생성 성능을 달성하는 효율적인 사전 학습 방법을 제시했습니다. 이러한 핵심 결과들은 135억 개의 매개변수를 가진 MoE 모델 여러 개를 2조 개의 토큰으로 학습하여 대규모로 검증되었습니다. 본 연구가 다중 모드 사전 학습을 이해하고 확장하는 데 필요한 체계적인 기반을 제공할 수 있기를 바랍니다.
Vision offers a critical axis for advancing foundation models, driving a shift towards natively unified multimodal pretraining. Despite this momentum, the design space and the fundamental mechanisms of how modalities interact during unified training remain underexplored. We provide empirical clarity through a systematic exploration of multimodal pretraining. Our controlled experiments on both synthetic and large-scale real-world datasets yield four key insights into the physics of multimodal pretraining: (i) Knowledge Flow: We disentangle how language, visual understanding, and visual generation transfer knowledge across modalities, revealing distinct patterns of influence and asymmetry; (ii) Synergy vs. Competition: We show that data "complexity" largely determines whether modalities are synergistic, identify architectural choices that promote synergy: such as shared attention and normalization with modality-specific feed-forward layers, and find that these behaviors generalize across different visual tokenizer designs; (iii) Early Unification: Unifying modalities from the very early stages and training them jointly is shown to be more effective than late alignment or sequential training. This process uncovers a vision laziness phenomenon, where delayed integration leads models to rely on language priors; (iv) Recipes: We derive efficient pretraining recipes that achieve strong generative performance using only 5% of the compute budget. These core findings are subsequently validated at scale by training multiple 13.5B MoE models on 2T tokens. We hope this study provides a principled foundation for understanding and scaling multimodal pretraining.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.