2602.03134v1 Feb 03, 2026 cs.CV

SwiftVLM: 크로스 레이어 토큰 바이패스를 통한 효율적인 비전-언어 모델 추론

SwiftVLM: Efficient Vision-Language Model Inference via Cross-Layer Token Bypass

Xin Miao
Xin Miao
Citations: 88
h-index: 5
Chen Qian
Chen Qian
Citations: 44
h-index: 2
Xinran Yu
Xinran Yu
Citations: 5
h-index: 2
Danyang Li
Danyang Li
Citations: 8
h-index: 2
Guoxuan Chi
Guoxuan Chi
Tsinghua University
Citations: 309
h-index: 7
Zheng Yang
Zheng Yang
Citations: 5
h-index: 2
Qiang Ma
Qiang Ma
Citations: 47
h-index: 3

비전-언어 모델(VLM)의 계산 비용을 줄이는 유망한 방법인 시각적 토큰 가지치기는 종종 효율성 향상을 위해 초기 단계의 가지치기 결정에 의존합니다. 이러한 방법은 거칠고 일반적인 추론 작업에서는 효과적이지만, 세밀한 시각적 디테일이 필요한 작업에서는 성능 저하가 심각합니다. 본 연구에서는 레이어별 분석을 통해, 레이어 간에 시각적 토큰의 중요도에 상당한 차이가 있음을 밝혀냈습니다. 얕은 레이어에서 중요하지 않은 것으로 판단되는 토큰이 텍스트 기반 추론에 매우 중요한 역할을 할 수 있습니다. 이러한 중요한 정보의 손실을 방지하기 위해, 우리는 '바이패스(bypass)'라는 새로운 가지치기 패러다임을 제안합니다. 이 패러다임은 선택되지 않은 시각적 토큰을 보존하고, 후속 가지치기 단계에서 재평가를 위해 전달합니다. 이를 바탕으로, 우리는 SwiftVLM을 제안합니다. SwiftVLM은 모델별 레이어에서 강력한 시각적 토큰 선택 기능을 활용하여 가지치기를 수행하는 간단하고 학습이 필요 없는 방법이며, 각 레이어에서 독립적인 가지치기 결정을 내릴 수 있도록 합니다. 다양한 VLM 및 벤치마크를 사용한 실험 결과, SwiftVLM은 기존의 가지치기 전략보다 우수한 성능을 보이며, 더 나은 정확도-효율성 균형과 더 신뢰할 수 있는 시각적 토큰 선택 동작을 제공하는 것으로 나타났습니다.

Original Abstract

Visual token pruning is a promising approach for reducing the computational cost of vision-language models (VLMs), and existing methods often rely on early pruning decisions to improve efficiency. While effective on coarse-grained reasoning tasks, they suffer from significant performance degradation on tasks requiring fine-grained visual details. Through layer-wise analysis, we reveal substantial discrepancies in visual token importance across layers, showing that tokens deemed unimportant at shallow layers can later become highly relevant for text-conditioned reasoning. To avoid irreversible critical information loss caused by premature pruning, we introduce a new pruning paradigm, termed bypass, which preserves unselected visual tokens and forwards them to subsequent pruning stages for re-evaluation. Building on this paradigm, we propose SwiftVLM, a simple and training-free method that performs pruning at model-specific layers with strong visual token selection capability, while enabling independent pruning decisions across layers. Experiments across multiple VLMs and benchmarks demonstrate that SwiftVLM consistently outperforms existing pruning strategies, achieving superior accuracy-efficiency trade-offs and more faithful visual token selection behavior.

1 Citations
0 Influential
3.5 Altmetric
18.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!