2607.28341v1 Jul 30, 2026 cs.CV

다중 모달 대규모 언어 모델에서 학습 없이 토큰을 제거하는 방법: 토큰 경향성 포착

Capturing Token Tendencies for Training-Free Token Pruning in Multimodal Large Language Models

Jiayi Ji
Jiayi Ji
Citations: 983
h-index: 16
Xiaoshuai Sun
Xiaoshuai Sun
Citations: 1,235
h-index: 20
Jie Ma
Jie Ma
Citations: 10
h-index: 2
Rongrong Ji
Rongrong Ji
Citations: 1,027
h-index: 17
Jie Gao
Jie Gao
Citations: 219
h-index: 3
Zhike Qiu
Zhike Qiu
Citations: 9
h-index: 2
Qianyu Chen
Qianyu Chen
Citations: 0
h-index: 0

다중 모달 대규모 언어 모델(MLLM)의 효율성을 높이기 위해 시각적 토큰 제거는 필수적이지만, 기존의 학습이 필요 없는 방법들은 치명적인 한계점을 가지고 있습니다. 이러한 방법들은 되돌릴 수 없는 필터링을 수행하기 위해 정적이고 순간적인 휴리스틱에 의존하며, 이는 MLLM의 계층 구조를 고려하지 않습니다. MLLM에서 토큰의 중요성은 종종 고정되어 있는 것이 아니라 동적으로 변화합니다. 결과적으로, 심층 레이어 추론에 중요한 토큰들이 얕은 레이어에서의 예측으로 인해 조기에 제거되는 경우가 많습니다. 이러한 문제를 해결하기 위해, 우리는 토큰의 경향성을 고려하는 새로운 프레임워크인 '트렌드 기반 제거(Trend-aware Pruning)'를 제안합니다. 기존 방법이 개별적인 점수에 의존하는 반면, 우리의 방법은 어텐션 흐름의 추세를 파악하여 동적으로 중요한 정보를 가진 토큰을 재활성화합니다. 이를 통해 초기에는 중요도가 낮게 평가되었지만 시간이 지남에 따라 의미적 중요성이 증가하는 '늦게 활성화되는' 토큰들을 선택적으로 복구하여 중요한 시각적 정보 손실을 방지합니다. 다양한 다중 모달 작업에서 수행한 실험 결과, 우리의 방법은 효율성과 성능 간의 균형이 뛰어나며, 전반적인 시각적 토큰 수를 77.8% 이상 줄이고 최종 레이어에 약 23개의 토큰만 남기면서도 경쟁력 있는 성능을 유지합니다. 이는 고효율 다중 모달 추론을 위한 강력하고 되돌릴 수 있는 솔루션을 제공합니다.

Original Abstract

While visual token pruning is essential for efficient Multimodal Large Language Models (MLLMs), existing training-free methods suffer from a critical limitation: they rely on static, instantaneous heuristics to perform irreversible filtering. This approach ignores the hierarchical nature of MLLMs, where token importance often evolves dynamically rather than remaining fixed across layers. Consequently, tokens essential for deep-layer reasoning are often prematurely discarded by shallow-layer estimates. To address this, we propose Trend-aware Pruning, a novel framework that elevates pruning from a local snapshot decision to a temporal trajectory modeling problem. Instead of relying on isolated scores, our method captures the momentum of attention flow. This enables a dynamic rectification mechanism that selectively reactivates "late-blooming" tokens, those initially undervalued but exhibiting rising semantic importance, thereby preventing the loss of critical visual cues. Extensive experiments demonstrate that our approach achieves a superior efficiency-performance trade-off across diverse multimodal tasks. Notably, it reduces visual tokens by over 77.8%, retaining only approximately 23 tokens in the final layer while maintaining competitive performance, offering a robust and reversible solution for high-efficiency multimodal inference.

0 Citations
0 Influential
10 Altmetric
50.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!