2607.11436v1 Jul 13, 2026 cs.AI

다중 모드 집중의 흐름: 시각적 정보 전달 기간 조정을 통한 지식 기반 VLM 추론

The Ebb and Flow of Multimodal Focus: Scheduling Visual Relay Windows for Grounded VLM Reasoning

H. Shen
H. Shen
Citations: 51
h-index: 2
Yi Bin
Yi Bin
Citations: 58
h-index: 4
Wen Ye
Wen Ye
Citations: 362
h-index: 5
Yujuan Ding
Yujuan Ding
Citations: 55
h-index: 3
Hongye Fang
Hongye Fang
Citations: 0
h-index: 0
Zheng Wang
Zheng Wang
Citations: 91
h-index: 3
Xing Xu
Xing Xu
Citations: 244
h-index: 7
Jingkuan Song
Jingkuan Song
Citations: 15,399
h-index: 63
Yun Zhang
Yun Zhang
Citations: 482
h-index: 4
Sirui Da
Sirui Da
Citations: 0
h-index: 0

시각-언어 모델은 다중 모드 추론 벤치마크에서 점점 더 높은 성과를 보이고 있지만, 시각적 증거가 언어 처리 단계에 진입하면 불안정해져 근거 기반 추론 능력이 약화되는 경향이 있습니다. 이러한 취약성을 이해하기 위해 우리는 VLM의 내부 작동 방식을 메커니즘적인 관점에서 분석하고, 깊이에 따른 다중 모드 주의 집중의 안정적인 세 단계를 발견했습니다. 이는 초기 질문 조건에 따른 정보 구성 단계, 중요한 중간 단계에서의 시각 중심 정보 전달 단계, 그리고 최종 답변 생성 단계로 구성됩니다. 우리는 중간 단계를 '시각적 정보 전달 기간(Visual Relay Window, VRW)'으로 정의하고, VRW의 형태가 작업 요구 사항에 따라 변하며, 근거 기반 생성과 인과적으로 관련되어 있으며, 뒷받침되지 않는 답변과 더 강력한 추론 과정을 구별하는 특징이라는 것을 보여줍니다. 이러한 내부 리듬을 바탕으로, 우리는 가벼운 학습 모듈을 사용하는 작업 적응형 추론 시간 제어 프레임워크인 TRACE를 제안합니다. TRACE는 사전 처리 단계에서 정보 전달 방식을 재조정하고, 디코딩 단계에서 수집된 시각적 정보를 유지합니다. 네 가지 공개 VLM 모델과 일곱 개의 벤치마크를 사용하여 실험한 결과, TRACE는 근거에 민감한 설정에서 평균적으로 4.33점, 최대 6.6점의 성능 향상을 보여주었으며, 추론 능력이 중요한 작업에서도 개선되었습니다. 이러한 결과는 깊이에 따른 다중 모드 집중을 명시적으로 제어하는 것이 근거 기반 다중 모드 추론 능력을 강화하는 효과적인 방법임을 보여줍니다.

Original Abstract

Vision-language models increasingly succeed on multimodal reasoning benchmarks, yet their visual evidence often becomes unstable once it enters the language stack, weakening evidence-grounded reasoning. To understand this fragility, we examine the internal dynamics of VLMs through a mechanistic lens and uncover a stable three-stage redistribution of multimodal attention focus across depth: an early question-conditioned organization, a critical middle visual-dominant relay, and a late return to answer formation. We operationalize the middle phase as the Visual Relay Window (VRW), and show that its geometry varies with task demand, is causally tied to grounded generation, and distinguishes unsupported answers from stronger reasoning trajectories. Guided by this internal rhythm, we propose TRACE, a task-adaptive inference-time control framework with lightweight trained modules. It reshapes relay allocation during prefill and preserves assembled visual support after handoff during decoding. Across four open-weight VLM backbones and seven benchmarks, TRACE delivers large gains on grounding-sensitive settings, improving them by 4.33 points on average and by up to 6.6 points, while also improving reasoning-heavy tasks. These results show that explicitly controlling multimodal focus across depth offers a unified and effective mechanism for strengthening evidence-grounded multimodal reasoning.

0 Citations
0 Influential
30 Altmetric
150.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!