2605.29657v1 May 28, 2026 cs.CV

OccamToken: 학습 없이 예산에 적응하는 토큰 제거를 통한 효율적인 VLM 추론

OccamToken: Efficient VLM Inference with Training-Free and Budget-Adaptive Token Pruning

Tuo An
Tuo An
Citations: 67
h-index: 4
Jianfei Yang
Jianfei Yang
Citations: 27
h-index: 2
Kuang Zuo
Kuang Zuo
Citations: 20
h-index: 1
Gen Li
Gen Li
Citations: 147
h-index: 2
Guohao Chen
Guohao Chen
Citations: 188
h-index: 6
Ting Chen
Ting Chen
Citations: 47
h-index: 2
Shilin Shan
Shilin Shan
Citations: 31
h-index: 4
Bofan Lyu
Bofan Lyu
Citations: 7
h-index: 1

비전-언어 모델(VLM)은 시각적 이해를 위해 긴 시각적 토큰 시퀀스에 의존하며, 이로 인해 사전 채움 단계에서 계산 및 메모리 비용이 많이 듭니다. 대부분의 기존 가지치기 방법은 절대 순위 방식을 따르며, 시각적 토큰에 중요도 점수를 할당하고 고정된 상위 K개 부분 집합을 유지합니다. 본 연구에서는 이러한 방식이 근본적으로 불안정하다고 주장합니다. 어텐션 싱크는 토큰의 중요도 순위를 왜곡하며, 이미지 중복과 쿼리에 의존적인 시각적 증거는 입력 데이터에 따른 고정된 토큰 예산의 신뢰성을 떨어뜨립니다. 우리는 학습이 필요 없는 프레임워크인 OccamToken을 제안합니다. OccamToken은 절대 토큰 순위를 대신하여 레지스터 기반 상대적 증거 테스트를 사용합니다. 기존 방식과 달리, 어떤 토큰이 전역적으로 중요한지를 묻는 것이 아니라, 시각적 토큰이 레지스터 기반 참조와 비교했을 때 추가적인 정보를 제공하는지 평가합니다. 핵심 아이디어는 레지스터 토큰이 본질적으로 낮은 정보 어텐션 패턴을 흡수하여 진정으로 유용한 시각적 증거를 식별하는 안정적인 참조 역할을 한다는 것입니다. 이러한 원칙에 따라, OccamToken은 레지스터 어텐션에서 파생된 동적 임계값을 사용하여 이미지 적응형 중복 제거 및 쿼리 적응형 관련성 제거를 동시에 수행합니다. LLaVA-NeXT, LLaVA-v1.5 및 Qwen3-VL 모델에 대해 OccamToken은 추가적인 학습 없이 정확도와 효율성의 균형을 지속적으로 향상시킵니다. 특히, LLaVA-NeXT에서 2880개의 시각적 토큰을 약 40개로 줄이면서 원래 정확도의 93% 이상을 유지하여 극단적인 1.4% 보존 비율에서도 안정적인 시각적 토큰 압축을 가능하게 합니다.

Original Abstract

Vision-language models (VLMs) rely on long visual token sequences for visual understanding, making the prefill stage expensive in both computation and memory. Most existing pruning methods follow an absolute-ranking paradigm, assigning importance scores to visual tokens and retaining a fixed top-K subset. In this work, we argue that this paradigm is fundamentally brittle: attention sinks distort token importance rankings, while image redundancy and query-dependent visual evidence make fixed token budgets unreliable across inputs. We propose OccamToken, a training-free framework that replaces absolute token ranking with register-anchored relative evidence testing. Instead of asking which tokens are globally important, OccamToken evaluates whether a visual token provides information beyond a register-based reference. Our key insight is that register tokens naturally absorb low-information attention patterns, making them a stable reference for identifying genuinely informative visual evidence. Based on this principle, OccamToken performs both image-adaptive redundancy pruning and query-adaptive relevance pruning through dynamic thresholds derived from register attention. Across LLaVA-NeXT, LLaVA-v1.5, and Qwen3-VL, OccamToken consistently improves the accuracy-efficiency trade-off without additional training. Notably, on LLaVA-NeXT, it reduces 2,880 visual tokens to approximately 40 while preserving over 93% of the original accuracy, enabling stable visual token compression even in the extreme 1.4% retention regime.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!