2607.12815v1 Jul 14, 2026 cs.AI

비전-언어 모델 추론에서의 시각 정보 접근 경계

Visual Access Boundaries in Vision-Language Model Reasoning

Gouki Minegishi
Gouki Minegishi
Citations: 136
h-index: 6
Shohei Taniguchi
Shohei Taniguchi
Citations: 71
h-index: 4
Masahiro Suzuki
Masahiro Suzuki
Citations: 50
h-index: 4
Yutaka Matsuo
Yutaka Matsuo
Citations: 78
h-index: 5
Hiroto Osaka
Hiroto Osaka
Citations: 0
h-index: 0
Kai Yamashita
Kai Yamashita
Citations: 0
h-index: 0

체인 오브 소트(Chain-of-Thought, CoT) 프롬프트는 비전-언어 모델(VLM)의 테스트 시간 확장 전략으로 널리 사용되지만, VLM이 더 긴 추론 과정을 생성할 때 실제로 무엇이 확장되는지는 명확하지 않습니다. 본 연구에서는 CoT가 이미지 토큰에 대한 지속적인 접근을 필요로 하는지, 아니면 주로 순방향 과정 초기에 제공된 시각 정보 위에서 작동하는지를 질문합니다. 우리는 Visual Access Sweep이라는 인과적 개입 방법을 도입하여, 생성된 토큰 쿼리가 레이어 깊이와 생성 시간에 따라 이미지 토큰 키에 대한 어텐션을 마스킹하고, 작업 정확도를 유지하는 최소 접근 영역을 Visual Access Boundary (VAB)라고 정의했습니다. Qwen2.5-VL 및 InternVL3의 여섯 가지 모델 구성에서, CoT 프롬프트를 사용하지 않은 직접 답변과 CoT 프롬프트 모두 유한한 VAB를 나타냅니다. Qwen2.5-VL-32B와 14B 및 38B 규모의 InternVL3에서, CoT가 CoT를 사용하지 않은 전체 접근 기준에 대해 평가될 때, 그 VAB 레이어는 상당하게 더 긴 생성 과정을 거치더라도 CoT를 사용하지 않은 경계와 최대 두 개의 레이어 차이를 보입니다. 이는 CoT가 추론 과정 전반에 걸쳐 직접적인 이미지 토큰 접근을 연장함으로써 성능을 향상시키는 것이 아니라, 이미지에서 파생된 숨겨진 상태 정보를 기반으로 언어 측면의 계산을 확장함으로써 성능을 향상시킨다는 것을 시사합니다. 또한, CoT의 효과는 인식적 출력(perceptual readout)에 의해 제한된다는 것을 보여줍니다. 쿼리되는 시각 속성을 모델이 안정적으로 추출할 수 있을 때 CoT가 도움이 되지만, 그렇지 않을 때는 도움이 되지 않습니다. 심볼릭 속성 오라클은 실제 속성이 텍스트로 제공될 경우 CoT가 계수를 개선할 수 있음을 보여주고, 단일 객체에 대한 프로브-vs-디코드 검증은 어려운 속성이 숨겨진 상태에서 선형적으로 복구 가능하지만 모델 자체는 이를 출력하기 어렵다는 것을 보여줍니다. 이러한 분석 결과들을 종합해 볼 때, 성능의 제한 요인은 계수(counting)가 아닌 출력(readout)에 있다는 결론을 내릴 수 있습니다.

Original Abstract

Chain-of-Thought (CoT) prompting is widely used as a test-time scaling strategy for Vision-Language Models (VLMs), but it remains unclear what is extended when VLMs generate longer reasoning traces. We ask whether CoT requires continued access to image tokens, or whether it mainly operates over visual information already made available earlier in the forward pass. We introduce Visual Access Sweep, a causal intervention that masks attention from generated-token queries to image-token keys along layer depth and generation time, and define the Visual Access Boundary (VAB) as the minimal access region that preserves task accuracy. Across six model configurations from Qwen2.5-VL and InternVL3, both no-CoT direct answering and CoT prompting exhibit finite VABs. In Qwen2.5-VL-32B and InternVL3 at 14B and 38B scales, when CoT is evaluated against the no-CoT full-access target, its VAB layer differs from the no-CoT boundary by at most two layers, despite substantially longer generations. This suggests that CoT does not primarily improve performance by prolonging direct image-token access throughout the reasoning trace, but by extending language-side computation over image-derived hidden-state information. We further show that CoT gains are constrained by perceptual readout. CoT helps when the queried visual attribute can be reliably read out by the model, but not when that readout is unreliable. A symbolic-attribute oracle shows that CoT can improve counting once ground-truth attributes are supplied as text, while a single-object probe-vs-decode check shows that hard attributes can be linearly recoverable from hidden states yet difficult for the model itself to output. Together, these analyses place the bottleneck at readout rather than counting.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!