사고 과정(Chain-of-Thought)이 다중 모드 LLM의 시각 공간 추론 능력을 저하시킨다
Chain-of-Thought Degrades Visual Spatial Reasoning Capabilities of Multimodal LLMs
사고 과정(Chain-of-Thought, CoT) 기반 사고를 활용하는 다중 모드 추론 모델(Multimodal Reasoning Models, MRM)은 수학 및 논리 문제 해결에 혁신을 가져왔습니다. 그러나 본 연구에서는 이러한 패러다임이 일반적인 공간 지능을 처리하는 데 어려움을 겪는다는 것을 보여줍니다. 우리는 13개의 공간 벤치마크를 사용하여 17개의 모델을 종합적으로 평가하고, 중요한 격차를 확인했습니다. 즉, CoT 프롬프트는 시각 공간 추론 성능을 지속적으로 저하시킨다는 것입니다. 또한, 새로운 No-Image++ ablation 방법을 통해, MRM과 CoT 프롬프트 기반의 MLMs가 심각한 단축 경로 학습(shortcut learning)을 수행하며, 이미지가 없는 경우에도 텍스트 기반의 선입견에서 시각적 세부 정보를 환각(hallucinate)한다는 것을 입증했습니다. 이러한 결과는 공간 작업에 대한 텍스트 기반 CoT의 효능에 의문을 제기하며, 시각 중심의 추론 패러다임의 필요성을 강조합니다.
Multimodal Reasoning Models (MRMs) leveraging Chain-of-Thought (CoT) based thinking have revolutionized mathematical and logical problem-solving. However, we show that this paradigm struggles with generalized spatial intelligence. We perform a comprehensive evaluation of seventeen models across thirteen spatial benchmarks and identify a critical gap: CoT prompting consistently degrades performance in visual spatial reasoning. Furthermore, through a novel No-Image++ ablation, we demonstrate that MRMs and CoT prompted MLMs suffer from severe shortcut learning, and hallucinate visual details from textual priors even when the image is absent. These findings challenge the efficacy of text-only CoT for spatial tasks and underscore the need for vision-centric reasoning paradigms.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.