2608.04496v1 Aug 05, 2026 cs.CV

DIVE: 효율적인 비전-언어 모델을 위한 동적 반복 시각 증거 구성

DIVE: Dynamic Iterative Visual Evidence Construction for Efficient Vision-Language Models

Wei He
Wei He
Citations: 724
h-index: 7
Zijie Wang
Zijie Wang
Citations: 0
h-index: 0
Cheng Zhong
Cheng Zhong
Citations: 28
h-index: 2
Xiao An
Xiao An
Citations: 36
h-index: 2
Jiepan Li
Jiepan Li
Citations: 470
h-index: 7
Guangyi Yang
Guangyi Yang
Citations: 999
h-index: 14

비전-언어 모델(VLMs)에서 시각 입력은 종종 텍스트보다 훨씬 긴 토큰 시퀀스로 인코딩되어, 시각 토큰이 효율적인 추론의 주요 병목 현상이 됩니다. 최근 많은 연구들이 이 병목 현상을 해결하기 위해 토큰 중요도를 평가하고 낮은 점수의 토큰을 제거하는 방법을 제시합니다. 그러나 단일 단계에서의 평가만으로는 충분하지 않는데, 이는 토큰의 프롬프트 관련 유용성이 이미 유지된 증거에 따라 달라지기 때문입니다. 이러한 통찰력을 바탕으로, 저희는 DIVE (Dynamic Iterative Visual Evidence Construction)라는 학습이 필요 없는 프레임워크를 제안합니다. DIVE는 시각 토큰 제거를 동적 증거 구성으로 재구성하며, 잔여량 기반 점수를 사용하여 가장 높은 점수의 토큰을 반복적으로 선택하고, 시각 및 프롬프트 잔여량을 업데이트하여 이미 설명된 증거의 영향을 줄이며, 남은 토큰을 다시 평가합니다. 이러한 선택-업데이트-재평가 과정을 통해 상호 보완적이고 프롬프트 관련성이 높은 증거 집합을 구축합니다. 8개의 이미지 이해 벤치마크를 사용한 실험 결과, DIVE는 다양한 토큰 수에 대해 일관적으로 성능을 유지하는 것으로 나타났습니다. 시각 토큰을 88.9% 줄임에도 불구하고, DIVE는 원본 모델의 평균 성능의 98.2%를 유지합니다. 코드 및 관련 자료는 다음 GitHub 주소에서 확인할 수 있습니다: https://github.com/Zhong-Chenchen/DIVE.git.

Original Abstract

Visual inputs in vision-language models (VLMs) are often encoded into substantially longer token sequences than text, making visual tokens a major bottleneck for efficient inference. Abundant recent methods address this bottleneck by scoring token importance and pruning low-scoring tokens in a single pass. However, one-shot scoring is insufficient because a token's prompt-relevant usefulness depends on the evidence already retained. Motivated by this insight, we introduce DIVE (Dynamic Iterative Visual Evidence Construction), a training-free framework that recasts visual-token pruning as dynamic evidence construction. DIVE repeatedly selects the remaining token with the highest residual-conditioned score, updates the visual and prompt residuals to discount the evidence already explained, and re-evaluates the remaining tokens. This select-update-re-evaluate process builds a retained set of complementary, prompt-relevant evidence. Experiments across eight image-understanding benchmarks show that DIVE consistently preserves performance across token budgets. With an 88.9% reduction in visual tokens, DIVE retains 98.2% of the uncompressed model's average performance. Code is available at https://github.com/Zhong-Chenchen/DIVE.git.

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!