2608.01979v1 Aug 03, 2026 cs.CV

ET-Prune: 시각 토큰 가지치기에서 질문 기반 증거 활용 동적 예산 할당 기법

ET-Prune: Evidence-Aware Dynamic Budgeting for Visual Token Pruning in Text-Rich MLLMs

Linghe Kong
Linghe Kong
Citations: 189
h-index: 8
Shaoqiu Zhang
Shaoqiu Zhang
Citations: 4
h-index: 1
Yulun Zhang
Yulun Zhang
Citations: 67
h-index: 5
Kai Liu
Kai Liu
Citations: 97
h-index: 5
Zizhong Ding
Zizhong Ding
Citations: 0
h-index: 0
Junxian Li
Junxian Li
Citations: 366
h-index: 9
Xiao Xiao
Xiao Xiao
Citations: 307
h-index: 2

시각 토큰 가지치기는 멀티모달 대규모 언어 모델(MLLM)의 추론 비용을 줄이지만, 고정된 토큰 비율은 텍스트가 풍부한 입력에 적합하지 않습니다. 광학 문자 인식(OCR) 중심 작업에서 중요한 증거는 질문에 의해 명시되는 작은 수, 레이블 또는 필드일 수 있습니다. 따라서 무분별한 가지치기는 이러한 증거를 제거하는 동시에 시각적으로 눈에 띄지만 관련 없는 영역을 유지할 수 있습니다. 본 논문에서는 학습이 필요 없는 프레임워크인 ET-Prune을 제안합니다. 이 방법은 가지치기를 증거 할당 문제로 간주하며, 디코더 측의 부분 쿼리-키 블록에서 질문에 조건화된 증거를 추출하고, 텍스트와 유사한 공간 영역을 보호합니다. 또한 증거의 불확실성과 밀도를 샘플별 토큰 최소값으로 변환합니다. 세 단계에 걸쳐 중간 레이어 이벤트가 이러한 예산을 향해 시퀀스를 이동시키며, 확산되거나 텍스트가 밀집된 증거에는 더 많은 토큰을 유지하고 집중된 증거는 더욱 적극적으로 가지치기합니다. 실험 결과, ET-Prune은 모든 구성에서 하나의 결정적 패스만 수행했을 때, 다른 가지치기 방법과 비교하여 성능이 우수하거나 동등한 수준을 보였으며, 약 절반의 토큰만을 사용했습니다. OCRBench-v2 데이터셋에서 ET-Prune은 Qwen3-VL-8B 모델에서 1.80%p, InternVL3.5-8B 모델에서 0.68%p 더 높은 성능을 보였으며, 약 절반의 시각 토큰을 유지했습니다. MMBench v1.1 데이터셋에서는 54.45%의 평균 시각 토큰 유지율로 0.8467의 순수 정확도를 달성하여, 일반 모델(Vanilla)의 0.8437보다 우수한 성능을 보였습니다. 이러한 결과는 텍스트가 풍부한 멀티모달 추론에서 증거 기반 동적 예산 할당이 품질과 비용 측면에서 유리한 trade-off를 제공한다는 것을 보여줍니다.

Original Abstract

Visual token pruning reduces the inference cost of multimodal large language models, but a fixed token ratio is poorly matched to text-rich inputs. In OCR-centric tasks, decisive evidence can be a small number, label, or field whose relevance is specified by the question; indiscriminate pruning can erase that evidence while retaining visually salient but irrelevant regions. We present ET-Prune, a training-free framework that casts pruning as evidence allocation. It derives question-conditioned evidence from a decoder-side partial query-key block, safeguards text-like spatial regions, and converts evidence uncertainty and density into a sample-specific token floor. Three progressive middle-layer events then move the sequence toward this budget, retaining more tokens for diffuse or text-dense evidence and pruning concentrated evidence more aggressively. At the observed point estimates from one deterministic pass per configuration, ET-Prune leads or ties among pruned methods in all six backbone-benchmark comparisons at roughly half tokens. On OCRBench-v2, it leads the strongest pruned baselines by 1.80 and 0.68 percentage points on Qwen3-VL-8B and InternVL3.5-8B, respectively, while retaining about half of the visual tokens; on MMBench v1.1, it reaches 0.8467 circular exact-matching accuracy versus 0.8437 for Vanilla at 54.45% average visual-token retention. These results show a favorable observed quality-cost trade-off for evidence-aware dynamic budgeting in text-rich multimodal inference.

0 Citations
0 Influential
4.5 Altmetric
22.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!