2606.09131v1 Jun 08, 2026 cs.AI

후반 레이어 통합만으로 충분하다: 시각적 과포화 상태에서 멀티모달 대규모 언어 모델을 위한 이중 경로 비전 토큰 라우팅

Late-Layer Fusion is Enough: Dual-Path Vision Token Routing for Multimodal Large Language Models under Visual Saturation

Siyuan Liu
Siyuan Liu
Citations: 25
h-index: 3
Jinyang Wu
Jinyang Wu
Citations: 0
h-index: 0

멀티모달 대규모 언어 모델(MLLM)은 일반적으로 단일 모드 텍스트 모델링에 설계된 깊고 대칭적인 트랜스포머 구조를 상속하며, 이미지와 텍스트 토큰 모두에 동일한 계산을 적용합니다. 이러한 설계는 중요한 양방향성 차이점을 간과합니다. 즉, 이미지와 텍스트 토큰은 정보 밀도, 중복성 및 필요한 추론 깊이에서 상당한 차이를 보입니다. LLaVA-1.5의 레이어별 분석을 통해 비전 토큰이 중간 레이어에서 과포화되는 경향이 있음을 확인했습니다. 구체적으로, 텍스트-이미지 어텐션은 0층에서 0.68인 값이 4층에서는 0.07으로 감소하고, 18층 이후에는 약 0.04로 안정화되는 반면, 텍스트 토큰은 여전히 깊이 있는 의미 처리로부터 이점을 얻습니다. 이러한 결과는 아키텍처의 대칭성과 깊이에 따른 비모드 진화 사이의 불일치를 시사하며, 이는 과도한 시각적 계산과 심층적인 작업별 적응 과정에서 지각 표현의 가능성 있는 왜곡을 초래할 수 있습니다. 이에 따라 우리는 효율적인 MLLM을 위한 양방향성 라우팅 프레임워크인 이중 경로 비전 토큰 라우팅(DPVR)을 제안합니다. DPVR의 핵심 구현체인 DPVR-LF (Late-Layer Fusion)는 과포화 지점에 있는 비전 토큰을 단일 레이어의 학습 가능한 사이드 브랜치로 분리하고, 이미지 위치를 건너뛰는 13개 레이어로 구성된 텍스트 전용 순방향 연산을 수행하며, 시각 및 텍스트 스트림을 최종 레이어에서만 다시 병합합니다. 약 3%의 학습 가능 파라미터로 DPVR-LF는 표준 벤치마크에서 경쟁력 있는 멀티모달 성능을 유지하면서 깊은 트랜스포머 스택에서의 시각적 계산량을 줄입니다. 이러한 결과는 비전 토큰이 모든 깊은 언어 모델 레이어를 통과해야 한다는 기존의 가정에 도전하며, LLaVA 스타일의 MLLM에서 강력한 지각 능력을 유지하는 데 단일 후반 융합 레이어가 충분할 수 있음을 시사합니다.

Original Abstract

Multimodal large language models (MLLMs) commonly inherit the deep, symmetric Transformer backbone designed for unimodal text modeling, and apply the same computation uniformly to image and language tokens. This design overlooks a key modality asymmetry: image and text tokens differ substantially in information density, redundancy, and required reasoning depth. Through a layer-wise analysis of LLaVA-1.5, we observe that vision tokens tend to saturate in the middle layers. Specifically, text-to-image attention decreases from 0.68 at layer 0 to 0.07 by layer 4, and stabilizes near 0.04 after layer 18, whereas text tokens continue to benefit from deep semantic processing. These findings suggest a mismatch between architectural symmetry and depth-asynchronous modality evolution, resulting in redundant visual computation and possible drift in perceptual representations during deep task-specific adaptation. Motivated by this, we propose Dual-Path Vision Token Routing (DPVR), a modality-asymmetric routing framework for efficient MLLMs. Its core instantiation, DPVR-LF (Late-Layer Fusion), routes vision tokens at the saturation point into a one-layer trainable side branch, runs a thirteen-layer text-only forward that skips image positions in the deep stack, and re-fuses the visual and textual streams only at the final layer. With approximately 3% trainable parameters, DPVR-LF preserves competitive multimodal performance on standard benchmarks while reducing visual computation in the deep Transformer stack. The results challenge the conventional assumption that vision tokens must traverse all deep language-model layers, and indicate that a single late fusion layer can be sufficient for maintaining strong perceptual competence in LLaVA-style MLLMs.

0 Citations
0 Influential
1.5 Altmetric
7.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!