2608.06270v1 Aug 06, 2026 cs.AI

시각적 도구 사용의 착시 현상: 이미지 활용 사고 방식에 대한 인과 관계 분석

The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images

Lai Wei
Lai Wei
Shanghai Jiao Tong University
Citations: 227
h-index: 7
Bo Peng
Bo Peng
Citations: 38
h-index: 3
Zhiheng Wang
Zhiheng Wang
Citations: 130
h-index: 5
Chaochao Lu
Chaochao Lu
Citations: 47
h-index: 1

"이미지 활용" 패러다임은 다중 모드 LLM에 작물(crop) 및 확대/축소와 같은 능동적인 시각 연산 기능을 제공합니다. 그러나 이러한 연산을 사용하는 모델들은 종종 직접 추론 방식보다 미미하거나 오히려 부정적인 성능 향상을 보이며, 토큰 비용은 훨씬 더 높습니다. 또한, 이러한 모델들은 관련 없는 영역을 반복적으로 작물 처리하는 경우가 많고, 직접 추론 방식으로 올바르게 답할 수 있는 질문에 대해서는 실패하기도 합니다. 본 연구에서는 제공된 시각 정보가 답변에 실제로 인과적인 영향을 미치는지 묻습니다. 이를 위해, 우리는 시각적 도구 사용을 관찰 기반 경로와 행동 유발 단축경로를 분리하는 인과 그래프로 표현했습니다. 그런 다음, 정책(도구 사용과 직접 추론 비교), 경로(실행 중인 모든 관찰 데이터 손상), 그리고 단계(특정 접두사 하에서 하나의 개별 관찰 데이터를 가상으로 대체)의 세 가지 수준에서 개입을 통해 이를 분석합니다. 단계 수준에서의 측정 지표인 "시각 증거 이득(Visual Evidence Gain)"은 각 반환된 관찰 데이터가 기여하는 정도를 분리합니다. 6개의 대표 모델과 5가지 세분화된 시각 인지 벤치마크를 사용하여 정책 오류와 두 가지 실패 모드를 발견했습니다. "관찰하지 않고 추론(Calling Without Looking)"의 경우, 반환된 관찰 데이터는 답변에 인과적인 영향을 미치지 않습니다. "관찰하지만 계획 없이(Looking Without Planning)"의 경우, 관찰 데이터는 유용하지만 호출 순서가 일관성이 없습니다. 경로 수준에서의 분석은 정책 수준에서의 정확도 향상분을 분해하여, 특정 모델에서만 성능 향상이 집중되어 있음을 보여줍니다. 우리는 이러한 불일치를 "시각적 도구 사용의 착시 현상"이라고 부릅니다. 즉, 전체적인 정확도 향상은 있지만, 시각적 도구 사용은 다양한 실행 환경에서 인과적으로 효과적이지 않습니다. 관련 코드는 https://github.com/OpenCausaLab/CauAudit 에서 확인할 수 있습니다.

Original Abstract

The "thinking-with-images" paradigm equips multimodal LLMs with active visual operations such as crop-and-zoom. However, models using these operations often achieve only marginal or negative gains over direct inference at substantially higher token cost. They may also repeatedly crop irrelevant regions and fail on questions that direct inference answers correctly. We ask whether the returned visual evidence causally affects the answer. To answer this question, we formulate visual tool-use as a causal graph that separates observation-mediated paths from action-induced shortcuts. We then audit it through interventions at the three levels: policy (comparing tool-use with direct inference), trajectory (corrupting all observations during rollout), and step (counterfactually replacing one individual observation under a fixed prefix). Our step-level estimand, Visual Evidence Gain, isolates the contribution of each returned observation. Across six representative models and five fine-grained perception benchmarks, we uncover policy miscalibration with two failure modes. In Calling Without Looking, returned observations have no causal effect on the answer. In Looking Without Planning, observations are informative but the call schedule is incoherent. A trajectory-level diagnostic decomposes the policy-level accuracy gain and shows that the gain is concentrated in a Calibrated minority. We term this discrepancy the illusion of visual tool-use: despite aggregate accuracy gains, visual tool-use is not causally effective across a broad range of rollouts. The code is available at https://github.com/OpenCausaLab/CauAudit.

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!