2604.02816v1 Apr 03, 2026 cs.CV

QAPruner: 양자화 인식 시각 토큰 가지치기 기법을 활용한 다중 모드 대규모 언어 모델

QAPruner: Quantization-Aware Vision Token Pruning for Multimodal Large Language Models

Yongtao Wang
Yongtao Wang
Citations: 905
h-index: 11
Xinhao Wang
Xinhao Wang
Citations: 207
h-index: 4
Z. Xia
Z. Xia
Citations: 0
h-index: 0
Zhiwei Lin
Zhiwei Lin
Citations: 13
h-index: 1
Zhe Li
Zhe Li
Citations: 213
h-index: 3

다중 모드 대규모 언어 모델(MLLM)은 뛰어난 추론 능력을 보여주지만, 높은 계산 및 메모리 비용으로 인해 자원이 제한된 환경에서의 활용에 어려움이 있습니다. 양자화(PTQ)와 시각 토큰 가지치기는 일반적인 압축 기술이지만, 일반적으로 독립적인 최적화 방식으로 취급됩니다. 본 논문에서는 이 두 기술이 밀접하게 연관되어 있음을 보여줍니다. 시맨틱 기반 토큰 가지치기를 PTQ로 최적화된 MLLM에 무작정 적용하면, 수치적 안정성에 중요한 활성화 값의 이상치를 제거하여 낮은 비트 환경(예: W4A4)에서 양자화 오류를 악화시킬 수 있습니다. 이러한 문제를 해결하기 위해, 양자화를 고려한 시각 토큰 가지치기 프레임워크를 제안합니다. 저희 방법은 시뮬레이션된 그룹별 양자화 오류와 이상치 강도를 결합한 경량 하이브리드 민감도 지표를 도입합니다. 이 지표와 표준 시맨틱 관련성 점수를 결합하여, 저희 방법은 의미적으로 유용하고 양자화에 강한 토큰을 유지합니다. 표준 LLaVA 아키텍처에 대한 실험 결과, 저희 방법이 기존 방식보다 일관되게 우수한 성능을 보였습니다. 시각 토큰의 12.5%만 유지하는 공격적인 가지치기 비율에서도, 저희 프레임워크는 정확도를 2.24% 향상시키고, 가지치기 없는 밀집 양자화보다 더 나은 성능을 보였습니다. 저희가 알고 있는 한, 이 연구는 정확한 저비트 MLLM 추론을 위해 시각 토큰 가지치기와 PTQ를 명시적으로 공동 최적화하는 첫 번째 방법입니다.

Original Abstract

Multimodal Large Language Models (MLLMs) have shown strong reasoning ability, but their high computational and memory costs hinder deployment in resource-constrained settings. While Post-Training Quantization (PTQ) and vision token pruning are standard compression techniques, they are usually treated as independent optimizations. In this paper, we show that these two techniques are strongly coupled: naively applying semantic-based token pruning to PTQ-optimized MLLMs can discard activation outliers that are important for numerical stability and thus worsen quantization errors in low-bit regimes (\textit{e.g.}, W4A4). To address this issue, we propose a quantization-aware vision token pruning framework. Our method introduces a lightweight hybrid sensitivity metric that combines simulated group-wise quantization error with outlier intensity. By combining this metric with standard semantic relevance scores, the method retains tokens that are both semantically informative and robust to quantization. Experiments on standard LLaVA architectures show that our method consistently outperforms naive integration baselines. At an aggressive pruning ratio that retains only 12.5\% of visual tokens, our framework improves accuracy by 2.24\% over the baseline and even surpasses dense quantization without pruning. To the best of our knowledge, this is the first method that explicitly co-optimizes vision token pruning and PTQ for accurate low-bit MLLM inference.

0 Citations
0 Influential
5.5 Altmetric
27.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!