2607.26608v1 Jul 29, 2026 cs.CV

이질적인 대규모 멀티모달 언어 모델(MLLM) 융합에서의 지식 전달 메커니즘 이해: 간단한 선형 접근 방식

Understanding Knowledge Transfer Mechanism in Heterogeneous MLLM Fusion: A Simple Linear Approach

Hong Xie
Hong Xie
Citations: 6
h-index: 2
Yinghao Hou
Yinghao Hou
Citations: 0
h-index: 0
Jiahe Fan
Jiahe Fan
Citations: 0
h-index: 0
Yuanhao Pu
Yuanhao Pu
Citations: 351
h-index: 4
Zongyuan Chen
Zongyuan Chen
Citations: 4
h-index: 1

훈련 없이 이질적인 다중 모드 대규모 언어 모델(MLLM)을 융합하는 것은 다양한 규모의 기능 이전 경로를 제공하지만, 통합 성능 향상이 실제 작은 모델이 무엇을 상속받는지 명확하게 보여주지는 않습니다. 기존 연구는 대부분 제한된 작업 집합 또는 전체 지표에 기반하여 설계되고 평가되었으며, 평가가 더 광범위한 작업 모음으로 확장됨에 따라 서로 다른 규모 간의 기능 이전 가능성에 대한 이해가 부족합니다. 이 질문을 조사하기 위해, 우리는 이질적인 융합 과정에서의 교차 규모 지식 전달을 분석하는 간단한 선형 탐색 방법인 Cross-Scale Directional Parameter Injection (CDPI)를 소개합니다. 이론적 분석 결과, 지식 전달의 선택성은 공유된 주입 방향에 대한 기능 의존적인 응답에 의해 결정되며, 2차 미분 효과는 실제 전달 범위를 제한하는 것으로 나타났습니다. Qwen3-VL 모델 쌍 네 개와 12개의 멀티모달 벤치마크를 사용하여 실험한 결과, 일관된 선택성 패턴이 확인되었습니다. 즉, 추론 능력, 특히 고수준 추론 능력이 향상되는 반면, 인식 성능은 원래 대상 모델과 거의 동일하게 유지됩니다. 구성 요소별 분석 결과, 고수준 추론 능력 향상은 주로 언어 모델에서 비롯되며, 비율 분석 결과 긍정적인 선택적 전달은 주로 작은 비율의 경우에 발생하는 것으로 나타났습니다. 이러한 연구 결과는 이질적인 MLLM 융합을 광범위한 기능 상속이 아닌, 제한된 환경 내에서의 언어 측면 추론 능력 전달로 재해석합니다.

Original Abstract

Training-free fusion of heterogeneous multimodal large language models (MLLMs) provides a direct route for cross-scale capability transfer, yet improvements in aggregate performance do not reveal what a smaller model actually inherits. Existing studies are largely designed and evaluated on limited task sets or aggregate metrics; as evaluation expands to broader task collections, whether different capabilities can transfer across scales remains poorly understood. To investigate this question, we introduce Cross-Scale Directional Parameter Injection (CDPI), a simple linear probe to analyze cross-scale knowledge transfer during heterogeneous fusion. A local theoretical analysis indicates that knowledge transfer selectivity is determined at first order by capability-dependent responses to a shared injection direction, while second-order curvature effects constrain the effective transfer regime. Across four Qwen3-VL model pairs and twelve multimodal benchmarks, our experiments reveal a consistent pattern of selectivity: gains concentrate on reasoning, particularly high-level reasoning, whereas perception performance remains close to that of the original target model. Component-wise ablations further show that high-level reasoning gains arise primarily from the language model, while ratio analysis finds that positive selective transfer occurs mainly in the small-ratio regime. These findings recast cross-scale heterogeneous MLLM fusion as selective language-side reasoning transfer within a narrow, low-interference regime, rather than broad capability inheritance.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!