CoRe 헤드를 활용한 다중 모드 LLM의 기능적 희소성에 대한 메커니즘적 통찰
Mechanistic Insights into Functional Sparsity in Multimodal LLMs via CoRe Heads
다중 모드 대규모 언어 모델(MLLM)은 복잡한 시각-언어 작업에서 놀라운 성능을 보이지만, 이러한 모델이 복잡하고 노이즈가 많은 환경에서 쿼리와 관련된 시각적 특징을 어떻게 추출하는지에 대한 메커니즘은 여전히 불분명합니다. 본 논문에서는 심층적인 해석 연구를 통해 MLLM 내에 존재하는 중요한 구조적 특성인 교차 모드 검색에서의 기능적 희소성을 밝혀냅니다. 'Retrieval Attention Mass (RAM)'라는 토큰 수준의 지표를 활용하여, 'Context-aware Retrieval (CoRe) 헤드'라고 불리는 매우 전문화된 어텐션 헤드의 하위 집합을 식별하고 분석합니다. 다양한 시각 영역과 모델 크기에서, 우리는 명확한 기능적 분화를 관찰했습니다. CoRe 헤드는 전용 정보 추출기로 작동하는 반면, 대부분의 다른 헤드는 더 넓은 문맥 영역에 걸쳐 어텐션을 분산시킵니다. 인과적 개입 실험을 통해 이러한 전문화된 헤드의 필요성을 더욱 입증합니다. 상위 5%의 CoRe 헤드를 제거하기만 해도 다중 모드 추론 성능이 크게 저하되는 반면, 낮은 순위의 헤드를 제거하면 영향이 미미합니다. 또한, 가속 실험은 CoRe 헤드의 유용성을 검증하며, 이 국소화된 희소성을 활용함으로써 추론 속도를 크게 향상시키면서도 강력한 작업 성능을 유지할 수 있음을 보여줍니다. 본 연구 결과는 MLLM 내의 기능적 희소성의 구조적 원리를 밝히고, 메커니즘적 해석에 대한 현재 이해를 심화하며, 향후 아키텍처 설계 및 모델 최적화를 위한 이론적 기반을 제공합니다.
While Multimodal Large Language Models (MLLMs) demonstrate remarkable proficiency on complex vision-language tasks, the mechanisms by which they extract query-relevant visual features from complex, noisy contexts remain opaque. In this paper, we present an in-depth interpretability study that uncovers a profound structural property within MLLMs: functional sparsity in cross-modal retrieval. Leveraging a token-level metric termed Retrieval Attention Mass (RAM), we identify and characterize a highly specialized subset of attention heads, referred to as Context-aware Retrieval (CoRe) heads. Across diverse visual domains and model scales, we observe a clear functional division: CoRe heads act as dedicated information extractors, while most other heads distribute attention over broader contextual regions. Causal interventions further demonstrate the necessity of these specialized heads. Ablating only the top 5% of CoRe heads causes significant degradation in multimodal reasoning performance, whereas ablating lower-ranked heads has minimal effect. Moreover, acceleration experiments validate the utility of CoRe heads, showing that leveraging this localized sparsity significantly accelerates inference while maintaining robust task performance. Our findings reveal a structural principle of functional sparsity within MLLMs, refining the current understanding of mechanistic interpretability and laying a theoretical foundation that can inspire future architecture design and model optimization.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.