2607.26596v1 Jul 29, 2026 cs.CV

분리된 시각 처리: 모달리티별 트랜스포머 대체 기반의 효율적인 다중 모드 적응

Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution

Zhengqi Wen
Zhengqi Wen
Citations: 2,630
h-index: 27
Mingkuan Feng
Mingkuan Feng
Citations: 138
h-index: 6
Jianhua Tao
Jianhua Tao
Citations: 148
h-index: 7

다중 모드 대규모 언어 모델(MLLM)은 통일된 트랜스포머 아키텍처 내에서 시각 및 텍스트 이해를 통합하여 놀라운 성능을 보여줍니다. 그러나 이러한 모델의 모든 매개변수를 시각 지침 조정에 사용하면 계산 비용이 많이 들고, 종종 불필요합니다. 왜냐하면 네트워크의 깊은 계층에서 시각적 토큰과 텍스트 토큰의 표현 요구 사항이 크게 다르기 때문입니다. 본 논문에서는 효율적인 학습 프레임워크인 분리된 시각 처리(DVP)를 제안합니다. DVP는 사전 학습된 LLM의 상위 디코더 계층을 가볍고 독립적으로 훈련 가능한 단일 트랜스포머 블록으로 대체하여, 이 블록은 오직 시각적 토큰 처리에 전용됩니다. 특히, 디코더 계층의 첫 번째 절반에서 공유 처리 과정을 거친 후, 시각적 토큰과 텍스트 토큰이 분리됩니다. 시각적 토큰은 새로 초기화된 단일 트랜스포머 블록을 통해 처리되는 반면, 텍스트 토큰은 원래의 동결된 디코더 계층을 통과합니다. 그런 다음 두 스트림은 언어 모델링 헤드 전에 연결됩니다. 학습 과정에서 단일 트랜스포머 블록만 업데이트되므로, 훈련 가능한 매개변수의 수가 크게 줄어듭니다. LLaVA-1.5 프레임워크에 대한 실험 결과는 DVP가 전체 매개변수의 일부분만 사용하여 MME, POPE 및 ChartQA 벤치마크에서 경쟁력 있는 성능을 달성한다는 것을 보여줍니다. 이는 MLLM의 시각적 표현이 분리되고 매개변수 효율적인 경로를 통해 효과적으로 학습될 수 있음을 시사합니다.

Original Abstract

Multimodal large language models (MLLMs) have demonstrated remarkable capabilities by integrating visual and textual understanding within a unified transformer architecture. However, fine-tuning all parameters of these models for visual instruction tuning is computationally expensive and often unnecessary, as the representation requirements for visual and textual tokens diverge significantly in the deeper layers of the network. In this paper, we propose Decoupled Visual Processing (DVP), an efficient training framework that replaces the upper decoder layers of a pretrained LLM with a lightweight, independently trainable single transformer block dedicated exclusively to visual token processing. Specifically, after shared processing through the first half of the decoder layers, visual and textual tokens are split: visual tokens are routed through a newly initialized single transformer block while textual tokens continue through the original frozen decoder layers. The two streams are then concatenated before the language modeling head. During training, only the single transformer block is updated, dramatically reducing the number of trainable parameters. Experiments on the LLaVA-1.5 framework demonstrate that DVP achieves competitive performance on MME, POPE, and ChartQA benchmarks while training only a fraction of the total parameters, suggesting that visual representations in MLLMs can be effectively learned through a decoupled, parameter-efficient pathway.

0 Citations
0 Influential
13.5 Altmetric
67.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!