2602.17871v1 Feb 19, 2026 cs.CV

비전-언어 모델의 세밀한 지식 역량 이해

Understanding the Fine-Grained Knowledge Capabilities of Vision-Language Models

D. Ghosh
D. Ghosh
Citations: 11
h-index: 1
Yuhui Zhang
Yuhui Zhang
Citations: 32
h-index: 3
Ludwig Schmidt
Ludwig Schmidt
Citations: 57
h-index: 3

비전-언어 모델(VLM)은 시각적 추론, 문서 이해, 멀티모달 대화에 이르는 광범위한 시각적 질의응답 벤치마크에서 상당한 진전을 이루었다. 이러한 향상은 다양한 기반 모델, 정렬 아키텍처 및 훈련 데이터를 바탕으로 구축된 광범위한 VLM에서 명백하게 나타난다. 그러나 최근 연구들에 따르면, 이러한 모델들은 세밀한 시각적 지식을 테스트하는 전통적인 이미지 분류 벤치마크에서는 뒤처지는 것으로 나타났다. 우리는 세밀한 분류 벤치마크에서 다수의 최신 VLM을 테스트하고, 세밀한 지식과 기타 비전 벤치마크 간의 괴리를 초래하는 잠재적 요인을 식별한다. 일련의 절제 실험을 통해, 더 나은 LLM을 사용하는 것은 모든 벤치마크 점수를 균등하게 향상시키는 반면, 더 나은 비전 인코더는 세밀한 분류 성능을 불균형적으로 향상시킨다는 것을 발견했다. 더 나아가, 우리는 사전 학습 단계 또한 세밀한 성능에 필수적이며, 특히 사전 학습 중에 언어 모델 가중치의 동결이 해제될 때 그러하다는 것을 발견했다. 이러한 통찰은 VLM의 세밀한 시각적 이해 및 비전 중심 역량을 강화할 수 있는 기반을 마련한다.

Original Abstract

Vision-language models (VLMs) have made substantial progress across a wide range of visual question answering benchmarks, spanning visual reasoning, document understanding, and multimodal dialogue. These improvements are evident in a wide range of VLMs built on a variety of base models, alignment architectures, and training data. However, recent works show that these models trail behind in traditional image classification benchmarks, which test fine-grained visual knowledge. We test a large number of recent VLMs on fine-grained classification benchmarks and identify potential factors in the disconnect between fine-grained knowledge and other vision benchmarks. Through a series of ablation experiments, we find that using a better LLM improves all benchmark scores equally, while a better vision encoder disproportionately improves fine-grained classification performance. Furthermore, we find that the pretraining stage is also vital to fine-grained performance, particularly when the language model weights are unfrozen during pretraining. These insights pave the way for enhancing fine-grained visual understanding and vision-centric capabilities in VLMs.

2 Citations
0 Influential
1.5 Altmetric
9.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!