2604.16060v1 Apr 17, 2026 cs.CV

사고 과정(Chain-of-Thought)이 다중 모드 LLM의 시각 공간 추론 능력을 저하시킨다

Chain-of-Thought Degrades Visual Spatial Reasoning Capabilities of Multimodal LLMs

Vineeth N. Balasubramanian
Vineeth N. Balasubramanian
Citations: 390
h-index: 8
Sai Srinivas Kancheti
Sai Srinivas Kancheti
Citations: 50
h-index: 2
Aditya Kanade
Aditya Kanade
Citations: 61
h-index: 3
Tanuja Ganu
Tanuja Ganu
Citations: 982
h-index: 13

사고 과정(Chain-of-Thought, CoT) 기반 사고를 활용하는 다중 모드 추론 모델(Multimodal Reasoning Models, MRM)은 수학 및 논리 문제 해결에 혁신을 가져왔습니다. 그러나 본 연구에서는 이러한 패러다임이 일반적인 공간 지능을 처리하는 데 어려움을 겪는다는 것을 보여줍니다. 우리는 13개의 공간 벤치마크를 사용하여 17개의 모델을 종합적으로 평가하고, 중요한 격차를 확인했습니다. 즉, CoT 프롬프트는 시각 공간 추론 성능을 지속적으로 저하시킨다는 것입니다. 또한, 새로운 No-Image++ ablation 방법을 통해, MRM과 CoT 프롬프트 기반의 MLMs가 심각한 단축 경로 학습(shortcut learning)을 수행하며, 이미지가 없는 경우에도 텍스트 기반의 선입견에서 시각적 세부 정보를 환각(hallucinate)한다는 것을 입증했습니다. 이러한 결과는 공간 작업에 대한 텍스트 기반 CoT의 효능에 의문을 제기하며, 시각 중심의 추론 패러다임의 필요성을 강조합니다.

Original Abstract

Multimodal Reasoning Models (MRMs) leveraging Chain-of-Thought (CoT) based thinking have revolutionized mathematical and logical problem-solving. However, we show that this paradigm struggles with generalized spatial intelligence. We perform a comprehensive evaluation of seventeen models across thirteen spatial benchmarks and identify a critical gap: CoT prompting consistently degrades performance in visual spatial reasoning. Furthermore, through a novel No-Image++ ablation, we demonstrate that MRMs and CoT prompted MLMs suffer from severe shortcut learning, and hallucinate visual details from textual priors even when the image is absent. These findings challenge the efficacy of text-only CoT for spatial tasks and underscore the need for vision-centric reasoning paradigms.

3 Citations
0 Influential
6.5 Altmetric
35.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!