2608.04132v1 Aug 04, 2026 cs.CV

RUTA: 속도-효율 최적화를 통한 체계적인 시각 토큰 할당

RUTA: Principled Visual Token Allocation via Rate-Utility Optimization

Zhihua Wang
Zhihua Wang
Citations: 27
h-index: 2
Balu Adsumilli
Balu Adsumilli
Citations: 41
h-index: 4
Yilin Wang
Yilin Wang
Citations: 3,176
h-index: 16
Jiangyu Zou
Jiangyu Zou
Citations: 1
h-index: 1
Kede Ma
Kede Ma
Citations: 50
h-index: 3
Xiaoyu Xu
Xiaoyu Xu
Citations: 2
h-index: 1

고해상도 이미지와 긴 동영상은 멀티모달 추론 및 정밀한 인식을 위해 비전-언어 모델에 풍부한 컨텍스트를 제공하지만, 결과적으로 생성되는 긴 시각 토큰 시퀀스는 대규모 언어 모델 측의 계산 비용과 메모리 사용량을 증가시킵니다. 기존의 시각 토큰 감소 기법은 종종 미리 정의된 속도로 작동하며, 최근에는 입력 데이터에 따라 토큰 수를 조정하기 위해 방법론별 학습된 임계값 또는 중요도 예측기를 사용합니다. 본 논문에서는 RUTA(Rate-Utility Token Allocation)라는 체계적인 속도-효율 토큰 할당 방법을 소개합니다. RUTA는 사전 LLM 감소를 수행하며, 각 이미지-쿼리 쌍에 대해 유지할 토큰과 할당량을 동시에 학습합니다. RUTA는 쿼리에 따라 조건을 부여한 후보 토큰을 구성하고, 각 후보에 대한 유지 확률을 예측합니다. 학습 과정에서 이러한 확률은 독립적인 베르누이 게이트를 매개변수화하며, 이들의 합은 각 쌍에 대한 토큰 수를 추정하는 미분 가능한 값을 제공합니다. 유지된 토큰은 앵커 역할을 하며, 유지되지 않은 토큰으로부터의 정보를 의미적 유사성과 공간적 근접성을 기준으로 집계합니다. RUTA는 하위 작업 손실과 예상되는 토큰 사용량을 균형 있게 고려하는 페널티가 적용된 속도-효율 목적 함수로 최적화됩니다. 5개의 벤치마크에 대한 평균 성능을 기준으로, LLaVA-NeXT-7B 및 Qwen3-VL-8B 모델의 전체 토큰 기준 성능과 비교했을 때, RUTA는 시각 토큰의 각각 $2.0%$와 $4.2%$만을 사용하면서 작업 성능의 $88.2%$와 $94.4%$를 유지합니다.

Original Abstract

High-resolution images and long videos provide vision-language models with rich context for multimodal reasoning and fine-grained perception, but the resulting long visual token sequences make large language model-side computation and memory costly. Existing visual token reducers often operate at prescribed rates, while recent methods adapt token counts across inputs using method-specific learned thresholds or importance predictors. We introduce RUTA, a principled Rate-Utility Token Allocation method that performs pre-LLM reduction by jointly learning which tokens to retain and how many to allocate to each image-query pair. RUTA constructs query-conditioned candidate tokens and predicts a retention probability for each candidate. During training, these probabilities parameterize independent Bernoulli gates, while their sum provides a differentiable training-time estimate of the token count for each pair. Retained tokens serve as anchors that aggregate information from non-retained tokens according to semantic affinity and spatial proximity. RUTA is optimized with a penalized rate-utility objective that balances downstream task loss against expected token usage. Averaged across five benchmarks and measured relative to each backbone's full-token baseline, RUTA uses only $2.0\%$ and $4.2\%$ of visual tokens while preserving $88.2\%$ and $94.4\%$ of task performance on LLaVA-NeXT-7B and Qwen3-VL-8B, respectively.

0 Citations
0 Influential
8 Altmetric
40.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!