2607.26326v1 Jul 28, 2026 cs.CV

보는 것인가, 아는 것인가? 다중 모드 대규모 언어 모델에서의 시각적 맥락 민감성

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models

Chengzu Li
Chengzu Li
University of Cambridge
Citations: 1,188
h-index: 11
Jiaang Li
Jiaang Li
University of Copenhagen
Citations: 267
h-index: 8
Zhaochong An
Zhaochong An
Citations: 63
h-index: 6
Serge J. Belongie
Serge J. Belongie
Citations: 32
h-index: 4
Yifei Yuan
Yifei Yuan
Citations: 83
h-index: 5
Xi Liu
Xi Liu
Citations: 0
h-index: 0
V´esteinn Snæbjarnarson
V´esteinn Snæbjarnarson
Citations: 21
h-index: 2

다중 모드 대규모 언어 모델(MLLM)은 사전 학습된 언어 모델의 풍부한 지식을 시각 정보와 통합하여 강력한 성능을 달성합니다. 그러나 이러한 모델들은 종종 시각 중심적인 작업에서 실패하는데, 특히 시각적 증거가 사전 학습된 지식과 충돌하는 경우 더욱 그렇습니다. 본 연구에서는 두 가지 진단 방법을 사용하여 이러한 실패 사례를 분석합니다: (1) 이미지 재구성을 통해 시각 정보의 활용 여부를 파악하고, (2) 다중 모드 맥락 민감성(모델이 시각적 맥락을 얼마나 따르는지)을 측정합니다. 후자를 평가하기 위해, 본 연구에서는 다섯 가지 큰 범주(공간-시간, 색상, 개수, 크기, 무게)를 포괄하는 벤치마크인 WhatIfVis를 도입했습니다. 이 벤치마크의 질문들은 이미지 또는 사전 지식을 통해 답변될 수 있습니다. 분석 결과, 다음과 같은 세 가지 주요 결과를 얻었습니다: (i) 거칠고 일반적인 시각적 증거는 유지됩니다. 이는 동결된 MLLM의 최종 레이어 이미지 토큰으로부터 이러한 속성들이 재구성될 수 있기 때문입니다. 따라서 이러한 속성에 대한 질문에 실패하는 경우는 시각 인코딩 자체의 문제라기보다는, 인식 이후 단계에서의 활용 문제를 나타냅니다. (ii) 명시적으로 시각적 증거를 사용하거나 무시하도록 지시하더라도, 사전 학습된 모델(WhatIfVis에 대한 지도 학습을 수행하지 않은 모델)은 불안정한 시각적 맥락 민감성을 보입니다. 지도 학습(SFT)은 이러한 제어 가능성을 향상시키고 다양한 도메인에서 일반화됩니다. 또한, 활성화 패칭을 통해 6가지 모델 모두에서 특정 아키텍처 깊이에서 시각과 사전 지식 간의 균형이 어떻게 조정되는지를 더욱 명확하게 파악할 수 있습니다. (iii) 시각과 사전 지식 간의 균형은 학습된 벡터를 통해 제어될 수 있습니다. 이 가이드(steering) 벡터를 적용하면, 의도적인 지시 없이도 기존 모델보다 더 나은 제어 가능성을 얻을 수 있습니다. 이러한 결과들은 문제 해결의 핵심이 어디에 있는지 보여줍니다. 즉, 본 연구에서 다루는 거칠고 일반적인 속성에 대해서는 MLLM이 시각적 증거를 인코딩하지만, 이를 얼마나 활용할지는 안정적으로 제어하지 못하는 것입니다.

Original Abstract

Multimodal Large Language Models (MLLMs) achieve strong performance by integrating visual inputs with the rich priors of pretrained language models. However, they often fail on vision-centric tasks, especially when visual evidence conflicts with pretrained knowledge. We explore these failures separately using two diagnostic paradigms: (1) probing whether visual information is available, via image reconstruction, and (2) measuring multimodal context sensitivity, the extent to which the model follows visual context versus the language prior. To support the second, we introduce the WhatIfVis, a benchmark spanning five coarse-grained dimensions (spatial-temporal, color, count, size, and weight) whose questions admit answers from either the image or the prior. Our analysis yields three findings: (i) Coarse-grained visual evidence is preserved, as these attributes can be reconstructed from the final-layer image tokens of frozen MLLMs. Failures on questions about these attributes therefore point to post-perceptual utilization, rather than to degraded visual encoding during perception. (ii) Even when explicitly instructed to use or ignore visual evidence, vanilla models (without supervised fine-tuning on the WhatIfVis) show unstable visual context sensitivity. Supervised fine-tuning (SFT) improves this controllability and generalizes across domains, and activation patching further localizes the vision-versus-prior trade-off at architecture-specific depths across all six models. (iii) The vision-versus-prior trade-off is controllable along a learned vector. Applying this steering vector, even without any intent instruction, improves controllability over the vanilla model. Together, these results relocate the bottleneck, indicating that for the coarse attributes we study, MLLMs encode the visual evidence but cannot reliably control their reliance on it.

0 Citations
0 Influential
5.5 Altmetric
27.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!