2606.17389v1 Jun 16, 2026 cs.CV

시각 정보는 거짓말을 하고, 일관성이 진실을 말한다: 비전-언어 모델에서 공간적 주의 집중과 신뢰성 분리

Visuals Lie, Consistency Speaks: Disentangling Spatial Attention from Reliability in Vision-Language Models

Shikhar Shiromani
Shikhar Shiromani
Citations: 28
h-index: 3
Logan Mann
Logan Mann
University of California, Santa Barbara
Citations: 2
h-index: 1
Yixuan Xia
Yixuan Xia
Citations: 19
h-index: 1
A. Saravanan
A. Saravanan
Citations: 0
h-index: 0
I. Dave
I. Dave
Citations: 573
h-index: 10
Saad M. Ismail
Saad M. Ismail
Citations: 2
h-index: 1
Emily Huang
Emily Huang
Citations: 10
h-index: 2
Ruizhe Li
Ruizhe Li
Citations: 12
h-index: 2
Kevin Zhu
Kevin Zhu
Citations: 0
h-index: 0
Algoverse AI Research
Algoverse AI Research
Citations: 0
h-index: 0

다중 모드 기반 모델은 점점 더 추론 에이전트로 활용되고 있으며, 따라서 모델의 환각 가능성을 파악하는 '신뢰성' 확보가 매우 중요합니다. 일반적인 직관으로 알려진 '주의-신뢰 가정'은 관련 영역에 대한 집중적인 주의 집중이 신뢰할 수 있는 답변을 나타내고, 분산된 주의는 혼란을 의미한다고 주장합니다. 본 연구에서는 최신 비전-언어 모델(VLM)의 신뢰성 신호에 대한 체계적인 분석인 VLM Reliability Probe (VRP)를 통해 이러한 주장에 도전합니다. 시각 인코더의 시선 방향을 정량화하기 위해 구조적 주의 집중 지표인 클러스터 수(C_k)와 공간 엔트로피(H_s)를 도입하고, 레이어별로 이 값들의 변화(Delta H_s)를 추적했습니다. 분석 결과 '상징적 분리' 현상이 나타났습니다. 모델들은 종종 초기 단계에서 시각 특징을 '고정'하지만, 이후 주의 집중이 분산되어 초기 인식과 최종 생성 결과가 단절됩니다. '접지 가설'과는 달리, 공간적 주의 집중은 정확도와 거의 상관관계가 없는 것으로 나타났습니다(R ≈ 0.001). 대신 신뢰성은 생성 과정의 역학 및 내부 상태 분포에 의해 결정되는 현상입니다. 샘플링된 추론 경로 간의 일치율인 '자기 일관성'이 진실 여부를 예측하는 가장 중요한 지표로 나타났습니다(R = 0.429). 인과적 개입을 통해 분석한 결과, LLaVA는 취약한 후반 단계 병목 현상에 예측을 고정시키는 반면, PaliGemma와 Qwen2-VL은 신뢰성을 전체적으로 분산시켜, 예시에서 ~50% 이상의 가장 중요한 레이어가 손상되어도 안정적인 성능을 유지하는 등 뚜렷한 구조적 차이를 보이는 것을 확인했습니다. 현재 VLM의 경우, 신뢰성 신호는 시각적 접지 정보와 관련이 적으며, 생성 과정의 역학 및 내부 상태 분석을 통해 더욱 정확하게 추론할 수 있습니다.

Original Abstract

Multimodal Foundation Models are increasingly used as reasoning agents, making reliability, knowing when a model may hallucinate, critical. A common intuition, which we call the Attention-Confidence Assumption, holds that reliability follows from "structural" visual perception: tight attention on relevant regions should signal a trustworthy answer, while scattered attention signals confusion. We challenge this through the VLM Reliability Probe (VRP), a systematic cross-family study of reliability signals in contemporary Vision-Language Models (VLMs). We introduce structural-attention metrics, cluster counts (C_k) and spatial entropy (H_s), to quantify the visual encoder's gaze, and track its evolution (Delta H_s) across layers. This reveals a "Symbolic Detachment": models often "Early Lock" visual features only to diffuse attention later, severing early perception from final generation. Contrary to the grounding hypothesis, we find a "Cluster Failure": spatial attention has near-zero correlation (R approx 0.001) with accuracy. Instead, reliability is a phenomenon of generation dynamics and internal-state distributions. Self-Consistency, the agreement rate across sampled reasoning paths, is the dominant predictor of truth (R = 0.429). Scaling causal interventions exposes a sharp architectural divergence: LLaVA locks its prediction in a fragile late-stage bottleneck, whereas PaliGemma and Qwen2-VL distribute reliability globally, staying resilient even when ~50% or more of their most predictive layer is destroyed. For current VLMs, reliability signals are detached from visual grounding maps and are best inferred from generation-time dynamics and hidden-state probes.

0 Citations
0 Influential
5 Altmetric
25.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!