2606.08970v1 Jun 08, 2026 cs.AI

비전-언어 모델 선택을 위한 효과적인 라우터

An Effective Router for Vision-Language Model Selection

Dianhui Chu
Dianhui Chu
Citations: 228
h-index: 7
Zhiying Tu
Zhiying Tu
Citations: 142
h-index: 6
Can Wang
Can Wang
Citations: 11
h-index: 2
Shengwei Wang
Shengwei Wang
Citations: 5
h-index: 2
Bolin Zhang
Bolin Zhang
Harbin Institute of Technology
Citations: 154
h-index: 6

다양한 성능과 자원 요구 사항을 가진 비전-언어 모델(VLM)이 광범위하게 사용됨에 따라, 수많은 VLM 후보군 중에서 가장 적합한 모델을 선택하는 것이 사용자에게 어려운 과제입니다. 기존 연구에서는 언어 모델의 성능 역설 현상을 밝히고 이를 해결하기 위한 라우팅 방법을 제시했습니다. 그러나 VLM 선택을 위한 라우터 개발은 여전히 중요하지만 어려운 문제입니다. 이는 주로 다음과 같은 문제에 직면합니다: 1) 특수 목적 데이터 부족, 2) 비효율적인 특징 표현, 3) 경직된 모델 공간 및 비용이 많이 드는 적응. 본 논문에서는 VLM 선택을 위한 다중 모드 데이터 세트를 구축했습니다. 이 데이터 세트는 32,626개의 고유한 이미지-텍스트 질의에 대한 7가지 주요 VLM의 출력 결과를 포함합니다. 그런 다음, VLM 선택을 위한 라우터인 ARMS를 제안합니다. ARMS는 입력 신호에 VLM 프로필을 결합하고, 간단하지만 효과적인 아키텍처를 사용하여 질의 및 VLM 기능의 표현력을 향상시킵니다. ARMS가 새로운 VLM에 더 잘 적응할 수 있도록 두 가지 확장된 훈련 전략인 점진적 훈련과 독립적 훈련을 제안합니다. in-distribution 및 out-of-distribution 테스트 세트에 대한 실험 결과는 ARMS의 효과를 입증합니다. 특히, 저희의 훈련 전략을 사용함으로써 ARMS(크기가 800M에 불과함)는 더 넓은 VLM 공간에 적응하고 GPT-4o와 같이 수백 배 더 큰 상용 모델보다 우수한 성능을 보입니다. 저희의 코드, 모델 및 데이터 세트는 익명 저장소에서 사용할 수 있습니다.

Original Abstract

Vision-language models (VLMs) with varying performance and resource requirements are widely deployed, making it difficult for users to select the most appropriate one among numerous VLM candidates. Existing work reveals the performance paradox phenomenon in language models and focuses on routing methods to solve it. However, developing a router for VLM selection is still a critical yet challenging problem, which primarily faces: 1) lack of specialized data, 2) ineffective feature representation, and 3) rigid model space and costly adaptation. In this paper, we construct a multimodal dataset for VLM selection, containing the outputs of seven mainstream VLMs on 32,626 unique image-text queries. We then propose ARMS, a router for VLM selection. ARMS enhances input signals with VLM profiles, employs a simple but effective architecture to improve representations of queries and VLM capabilities. To improve ARMS' adaptation to new VLMs, we propose two extension training strategies: incremental training and independent training. Experimental results on both in-distribution and out-of-distribution test sets demonstrate the effectiveness of ARMS. In particular, using our training strategy, ARMs (only 800M in size) can adapt to a broader VLM space and defeat commercial models like GPT-4o that are hundreds of times larger in scale. Our code, models, and datasets are available in the anonymous repository.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!