TOPS: 시각 토큰 최적 보존 집합 구축을 통한 첫 번째 원리 기반 시각 토큰 가지치기를 활용한 효율적인 멀티모달 대규모 언어 모델 추론
TOPS: First-Principles Visual Token Pruning via Constructing Token Optimal Preservation Sets for Efficient MLLM Inference
멀티모달 대규모 언어 모델(MLLM)은 뛰어난 멀티모달 추론 능력을 보여주지만, 많은 수의 시각 토큰으로 인해 발생하는 상당한 계산 오버헤드로 인해 효율성이 제한됩니다. 시각 토큰 가지치기는 자연스러운 해결책을 제공하지만, 기존 방법들은 완벽하지 않습니다. 어텐션 기반 기준은 종종 불필요한 토큰을 유지하고, 다양성 기반 기준은 사용자 지시에 무관한 경우가 많습니다. 여러 기준을 결합하는 방법조차도 토큰 가지치의 근본적인 목표에 대한 체계적인 정의가 부족합니다. 본 논문에서는 시각 토큰 가지치기를 첫 번째 원리 관점에서 재검토하고, 이를 '시각 토큰 최적 보존 집합' 구축으로 공식화합니다. 상향식 정보 이론 분석을 통해 효과적인 토큰 선택을 위한 세 가지 핵심 원칙인 '작업 관련성', '정보 커버리지', 그리고 '의미적 다양성'을 식별했습니다. 이러한 원칙에 기반하여, 본 논문에서는 훈련이 필요 없고 모델에 독립적인 가지치기 모듈인 TOPS를 제안합니다. 다양한 MLLM 아키텍처와 14개의 벤치마크에서 수행한 광범위한 실험 결과는 TOPS가 다양한 가지치기 설정에서 기존 방법보다 우수한 성능을 보임을 입증합니다. 특히, LLaVA-NeXT 모델에서 TOPS는 시각 토큰의 77.8%를 제거하면서 7B 및 13B 모델 모두에서 100.0%와 100.6%의 성능을 유지했습니다. 이는 불필요한 시각 토큰을 제거하는 것이 환각 현상을 완화하고 경량 MLLM 설계에 영감을 줄 수 있음을 시사합니다.
Multimodal large language models (MLLMs) have achieved strong multimodal reasoning capabilities, but their efficiency is limited by the large number of visual tokens, which introduces substantial computational overhead. Visual token pruning offers a natural solution, yet existing methods are imperfect: attention-based criteria tend to retain redundant tokens, while diversity-based criteria are often agnostic to user instructions. Even methods that combine multiple criteria still lack a principled formulation of the intrinsic objective of token pruning. In this paper, we revisit visual token pruning from a first-principles perspective and formulate it as constructing Token Optimal Preservation Sets. Through a top-down information-theoretic analysis, we identify three fundamental principles for effective token selection: Task Relevance, Information Coverage, and Semantic Diversity. Based on these principles, we propose TOPS, a training-free and model-agnostic pruning module that can be applied to various MLLMs. Extensive experiments on 7 MLLM backbones and 14 benchmarks demonstrate that TOPS outperforms prior methods under diverse pruning settings. Notably, on LLaVA-NeXT, TOPS removes 77.8% of visual tokens while preserving 100.0% and 100.6% performance on its 7B and 13B models, respectively, suggesting that pruning redundant visual tokens can sometimes mitigate hallucination and inspire future lightweight MLLM design.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.