See2Think: 다중 모드 모델이 정말로 중간 시각 상태를 사용하는가?
See2Think: Do Multimodal Models Really Use Intermediate Visual States?
다중 모드 대규모 언어 모델은 추론 과정에서 스케치, 주석, 도구 및 중간 이미지를 점점 더 많이 사용하지만, 이러한 모델이 실제로 이러한 시각적 상태에 얼마나 의존하는지는 여전히 불분명합니다. 기존의 벤치마크는 작업 유형의 제한적인 범위 또는 부분적으로 텍스트로 해결 가능한 샘플과, 중간 시각 상태가 어떻게 생성되고 표현되며 사용되는지 진단하지 않고 최종 답변만을 강조하는 평가 방법으로 인해 한계점을 가지고 있습니다. 본 연구에서는 See2ThinkBench와 Visual Action-of-Thought (VAoT)를 포함하는 통합적인 평가 프레임워크인 See2Think를 소개합니다. See2ThinkBench는 2차원 구조, 3차원 장면 및 실제 세계 추론을 포괄하는 12가지 작업 유형에 걸쳐 1,200개의 개방형, 시각적으로 의존적인 문제를 포함하고 있습니다. VAoT는 네 가지 제어된 추론 환경에서 텍스트 기반 사고 과정, 시각적 동작, 표현된 상태 및 후속 추론 과정을 기록합니다. 대표적인 독점 및 오픈 소스 다중 모드 모델을 평가한 결과, 시각적 추론은 모델과 환경에 크게 의존하며, 어떤 특정 설정이 모든 작업 유형에서 일관적으로 우수한 성능을 보이는 것은 아니었습니다. 추가 분석 결과, 모델은 일반적으로 관련 시각적 작업을 선택하지만, 정확하고 신뢰할 수 있는 표현(rendering)이 가장 큰 걸림돌이며, 높은 피드백 활용도가 반드시 정확도 향상으로 이어지지는 않습니다. 작업과 관련된 잘못된 피드백 환경에서 모델은 시각적 상태에 대한 행동 의존성을 보이는 것으로 나타났으며, 제어된 개입 시 정확도는 10% 이상 감소했습니다.
Multimodal large language models increasingly use sketches, annotations, tools, and intermediate images during reasoning, but it remains unclear whether they truly rely on these visual states. Existing benchmarks are limited both by task collections with narrow coverage or partially text-solvable samples and by evaluations that emphasize final answers without diagnosing how intermediate visual states are generated, rendered, and used. We introduce See2Think, a unified evaluation framework comprising See2ThinkBench and Visual Action-of-Thought (VAoT). See2ThinkBench contains 1,200 open-ended, visually dependent problems across 12 task categories spanning 2D structured, 3D scene, and real-world reasoning. VAoT records textual thoughts, visual actions, rendered states, and subsequent reasoning under four controlled inference settings. Evaluating representative proprietary and open-source multimodal models, we find that visual reasoning is strongly model- and environment-dependent, with no single setting consistently dominating across tasks. Process analysis further shows that models usually select relevant visual operations, while faithful rendering remains the clearest bottleneck and high feedback uptake does not necessarily translate into accuracy gains. Under task-relevant corrupted feedback, models exhibit behavioral dependence on visual states, with accuracy dropping by over 10 percentage points in controlled interventions.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.