2607.27830v1 Jul 30, 2026 cs.CV

한 번의 사고로 충분하다: 고해상도 시각 질의응답을 위한 중간 레이어 증거 경로 설정

Thinking Once Is Enough: Intermediate-Layer Evidence Routing for High-Resolution VQA

Fei Shen
Fei Shen
Citations: 125
h-index: 6
Yong Dai
Yong Dai
Citations: 56
h-index: 3
Z. Mao
Z. Mao
Citations: 1
h-index: 1
Xianjie Liu
Xianjie Liu
Citations: 8
h-index: 1
Tianyu Meng
Tianyu Meng
Citations: 0
h-index: 0
Yidong Wang
Yidong Wang
Citations: 0
h-index: 0
Wenzhuo Zhao
Wenzhuo Zhao
Citations: 0
h-index: 0
Ronghao Xian
Ronghao Xian
Citations: 91
h-index: 2
Yao Jiang
Yao Jiang
Citations: 306
h-index: 7
Junfeng Fang
Junfeng Fang
Citations: 200
h-index: 6
Yi Zhang
Yi Zhang
Citations: 65
h-index: 3
Keren Fu
Keren Fu
Citations: 202
h-index: 9

고해상도 시각 질의응답(HR-VQA)은 종종 불충분한 정보 획득 문제로 간주되며, 멀티모달 대규모 언어 모델이 이미지를 다시 검사하기 위해 잘라내기, 재인코딩 또는 다단계 검색을 수행해야 합니다. 본 논문에서는 이러한 관점이 불완전함을 보여줍니다. 많은 경우, 세밀한 정보는 시각적 인코딩 과정에서 살아남아 중간 레이어 경로 설정 영역 내에서 식별 가능하고 중요한 역할을 하지만, 이후 답변 생성 단계에서 희석됩니다. 우리는 Thinking-Once라는 방법을 제안합니다. 이는 훈련 없이 단일 시각적 입력만으로 증거를 경로 설정하는 방법입니다. 이 방법은 질문에 따른 주의 메커니즘을 해당 영역에서 재구성하고, 핵심 개체 토큰과 간결한 배경 정보를 유지하며, 추가적인 시각적 인코딩 없이 이러한 정보를 후속 레이어로 전달합니다. 다섯 가지 기본 모델에서 Thinking-Once는 일관되게 성능 향상을 보였거나 기존 성능에 준하는 결과를 보여주었습니다. V$^*$Bench, HRBench-4K 및 HRBench-8K의 평균 점수를 각각 +3.1, +3.0, +2.7점 증가시키면서 평균 피크 메모리를 약 4GB 줄였습니다. Qwen2.5-VL-7B 모델에서 Thinking-Once는 세 가지 벤치마크에서 각각 +9.9, +4.6, +5.5점을 향상시켜 전체 벤치마크 평균을 72.5점에서 79.1점으로 끌어올렸습니다. ZwZ-8B 기본 모델에서는 Thinking-Once를 통해 평균 점수 82.7에 도달했습니다. 11개의 오픈 소스 HR-VQA 기준 모델과 비교했을 때, Thinking-Once는 세 가지 벤치마크 평균 모두에서 최고 또는 공동 최고 점수를 기록했으며, 전체 평균에서도 가장 높은 성능을 보였습니다. 예를 들어, DeepScan과 비교했을 때 V$^*$Bench 추론 시간을 97.2% 단축하면서 전체 벤치마크 평균을 77.8점에서 79.1점으로 향상시켰습니다. 이러한 결과는 HR-VQA의 성능을 향상시키기 위해서는 반복적으로 새로운 시각 정보를 획득하는 것보다 이미 인코딩된 정보를 효율적으로 활용하는 것이 중요하다는 것을 보여줍니다. 코드 내용은 부록에 포함되어 있습니다.

Original Abstract

High-resolution visual question answering (HR-VQA) is often treated as a problem of insufficient evidence acquisition, where failing multimodal large language models must inspect images again through cropping, re-encoding, or multi-round search. We show that this view is incomplete: in many cases, fine-grained evidence has already survived visual encoding and become identifiable and influential within an intermediate-layer routing window, but is later diluted before answer generation. We propose Thinking-Once, a \textbf{training-free, single-visual-pass} evidence-routing method that reconstructs question-conditioned attention at this window, preserves core entity tokens and compact background context, and routes this evidence to later layers without extra visual encoding. Across five base models, Thinking-Once consistently improves or matches the corresponding base setting, increasing the average scores on V$^*$Bench, HRBench-4K, and HRBench-8K by \textit{+3.1}, \textit{+3.0}, and \textit{+2.7} points while reducing the average peak memory by about 4,GB. On Qwen2.5-VL-7B, it improves the three benchmarks by \textit{+9.9}, \textit{+4.6}, and \textit{+5.5} points, raising the cross-benchmark mean from 72.5 to 79.1. With the ZwZ-8B base model, Thinking-Once reaches a mean score of 82.7. Against 11 open-source HR-VQA baselines, it obtains the best or tied-best score on all three benchmark averages and the best overall mean; for example, compared with DeepScan, it reduces V$^*$Bench inference time by \textbf{97.2\%} while improving the cross-benchmark mean from 77.8 to 79.1. These results show that HR-VQA can be improved by routing already encoded evidence rather than repeatedly acquiring new visual inputs. Code is available in the appendix.

0 Citations
0 Influential
4.5 Altmetric
22.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!