2605.28115v1 May 27, 2026 cs.AI

CIVIC: 효율적인 비전-언어 모델을 위한 엔드투엔드 시퀀스 압축

CIVIC: End-to-End Sequence Compactness for Efficient Vision-Language Models

Bo Yu
Bo Yu
Citations: 47
h-index: 3
Fengze Yang
Fengze Yang
Citations: 41
h-index: 4
Xuewen Luo
Xuewen Luo
Citations: 36
h-index: 3
Chenxi Liu
Chenxi Liu
Citations: 24
h-index: 2
Cathy Liu
Cathy Liu
Citations: 0
h-index: 0

비전-언어 모델(VLMs)은 고해상도 시각적 토큰으로 인해 심각한 메모리 및 지연 시간 병목 현상을 겪습니다. 현재의 토큰 감소 방법은 이론적으로 FLOPs를 절약하지만, 사후 가지치기(post-hoc pruning)는 구조적인 오버헤드를 발생시켜 실제 속도 향상에는 미미합니다. 반면, 연속적인 압축 경로를 강제하면 기하학적 왜곡 및 세밀한 위치 정보 손실의 위험이 있습니다. 이러한 문제점을 해결하기 위해, 본 논문에서는 경로 일관성을 유지하는 압축 시각 추론 프레임워크인 CIVIC을 제안합니다. CIVIC은 비전 인코더, 투영 레이어, LLM 사전 학습(prefill) 및 KV-캐시 전반에 걸쳐 압축된 시퀀스 표현을 원활하게 유지함으로써 불연속적인 메모리 접근 및 지역적 병합 오버헤드를 방지합니다. Qwen3-VL 아키텍처에서 CIVIC을 평가한 결과, 시퀀스 감소를 실제 물리적 하드웨어 효율성으로 성공적으로 변환하여 KV-캐시 메모리를 기본 모델의 약 1/3 수준으로 줄이고 엔드투엔드 추론 지연 시간을 단축했습니다. 본 논문에서는 텍스트 정렬된 KL 증류(KL distillation) 및 적응적인 공간 유지 비율을 활용하여, 엄격한 다중 모달 추론 및 시각적 위치 정보 인식 벤치마크에서 정확도를 저하시키지 않고 이러한 효율성 향상을 달성했습니다.

Original Abstract

Vision-Language Models (VLMs) face severe memory and latency bottlenecks due to high-resolution visual tokens. While current token reduction methods theoretically save FLOPs, post-hoc pruning introduces structural overhead, failing to yield proportional wall-clock acceleration. However, enforcing a contiguous compact pathway risks geometric disorientation and loss of fine-grained localization. To overcome these barriers, this paper introduces CIVIC, a path-consistent compact visual inference framework. By maintaining compact sequence representations seamlessly across the vision encoder, projection layer, LLM prefill, and KV-cache, CIVIC avoids non-contiguous memory access and localized unmerging overheads. Evaluated on the Qwen3-VL architecture, CIVIC successfully translates sequence reductions into genuine physical hardware efficiency, shrinking KV-cache memory to approximately one-third of the baseline and reducing end-to-end inference latency. Enabled by text-aligned KL distillation and an adaptive spatial retention floor, CIVIC achieves these efficiency milestones without degrading accuracy across rigorous multimodal reasoning and visual grounding benchmarks.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!