2606.16122v1 Jun 15, 2026 cs.AI

시각적 근거를 활용한 사고

Thinking with Visual Grounding

Yihe Deng
Yihe Deng
Citations: 407
h-index: 7
Junkai Zhang
Junkai Zhang
University of California, Los Angeles
Citations: 224
h-index: 10
Wei Wang
Wei Wang
Citations: 382
h-index: 7
Kai-Wei Chang
Kai-Wei Chang
Citations: 990
h-index: 4

시각적 사고는 단순히 논리적으로 타당할 뿐만 아니라, 그 증거를 보여주어야 합니다. 최근의 비전-언어 모델(VLMs)은 자연어 추론 과정을 생성할 수 있지만, 이러한 과정에서 사용되는 시각적 정보가 명시적으로 드러나지 않아 검증하기 어렵고 학습시키기 어렵습니다. 본 연구에서는 '시각적 근거를 활용한 사고'라는 새로운 접근 방식을 제안합니다. 이는 모델이 자연어 추론 과정과 함께 각 단계에서 사용되는 시각적 증거에 대한 명시적인 점 또는 사각형 영역 정보를 동시에 제공하는 방식입니다. 이를 통해 모델은 중간 추론 과정을 언어로 표현하면서, 관련된 이미지 영역의 핵심 객체를 정확하게 연결할 수 있습니다. 이러한 기능을 학습하기 위해, 우리는 올바른 시각적 추론 과정을 생성하고, 해당 과정에서 필요한 시각적 객체를 추출하며, SAM3 기반 에이전트를 사용하여 이를 명시적으로 연결하는 확장 가능한 합성 파이프라인을 구축했습니다. 또한, 생성된 객체 참조가 실제 이미지 증거와 일치하는지 평가하여 보상하는 '근거 인식 강화 학습' 방법을 제안합니다. 두 가지 개수 세기 벤치마크와 네 가지 공간 추론 벤치마크에서, Gemma3-4B-IT 모델에 시각적 근거를 활용한 사고 방식을 적용하면 원래 모델과 단순히 언어적 추론만 수행하는 모델보다 성능이 일관되게 향상됩니다. 특히 공간 추론 분야에서는, 시각적 근거를 활용한 4B 모델이 동일 모델 패밀리의 더 큰 Gemma3-27B-IT 모델과 동등하거나 오히려 더 나은 성능을 보이는 경우도 있었습니다. 분석 결과, 점 영역 정보는 개수 세기에 적합하며, 사각형 영역 정보는 명시적인 근거 보상을 통해 공간 추론 작업에서 가장 큰 효과를 나타내는 것으로 확인되었습니다. 전반적으로, 본 연구의 결과는 VLMs가 언어적 사고 과정과 함께 해당 과정을 뒷받침하는 이미지 영역 정보를 제공할 때 더욱 정확하고 효율적인 사고를 수행할 수 있음을 보여줍니다.

Original Abstract

Visual thinking should not only sound right; it should show its evidence. While recent vision-language models (VLMs) can produce natural-language reasoning traces, these traces often leave the supporting image regions implicit, making them hard to verify and difficult to supervise. We introduce visually grounded thinking, a reasoning process in which models interleave natural-language thoughts with explicit point or box groundings of the visual evidence used at each step. This lets the model express intermediate reasoning in language while grounding key objects in the image regions they refer to. To train this behavior, we construct a scalable synthesis pipeline that distills correct visual reasoning traces, extracts the visual objects required by the traces, grounds them with a SAM3-based agent, and derives aligned point and box supervision from the resulting masks. We further propose grounding-aware reinforcement learning, which combines answer correctness rewards with dense grounding rewards that score whether generated object references match the correct image evidence. Across two counting benchmarks and four spatial reasoning benchmarks, adding visually grounded thinking to Gemma3-4B-IT consistently improves performance over the original model and the non-grounded thinking baseline. On spatial reasoning, the visually grounded thinking 4B models match, and in some cases surpass, Gemma3-27B-IT from the same model family. Our analysis shows that point grounding is well suited to counting, while box grounding benefits most from explicit grounding rewards on spatial tasks. Overall, our results show that VLMs think better when their intermediate thoughts are tied to the image regions that make them true.

0 Citations
0 Influential
5 Altmetric
25.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!