2602.02465v1 Feb 02, 2026 cs.AI

MentisOculi: 심상을 활용한 추론의 한계 규명

MentisOculi: Revealing the Limits of Reasoning with Mental Imagery

Thaddäus Wiedemer
Thaddäus Wiedemer
Citations: 338
h-index: 6
Thomas Klein
Thomas Klein
Citations: 47
h-index: 3
Prasanna Mayilvahanan
Prasanna Mayilvahanan
Citations: 87
h-index: 4
Matthias Bethge
Matthias Bethge
Citations: 53
h-index: 3
F. Wichmann
F. Wichmann
Citations: 17
h-index: 2
W. Brendel
W. Brendel
Citations: 12
h-index: 2
J. Zeller
J. Zeller
Citations: 94
h-index: 6
Fanfei Li
Fanfei Li
Citations: 6
h-index: 1
R. Cotterell
R. Cotterell
Citations: 90
h-index: 3

프론티어 모델들은 단순히 시각 정보를 입력받는 멀티모달 대형 언어 모델(MLLM)에서 네이티브 인터리브 생성이 가능한 통합 멀티모달 모델(UMM)로 전환되고 있습니다. 이러한 변화는 인간의 심상과 유사하게 중간 단계의 시각화를 추론 보조 도구로 사용하는 것에 대한 관심을 불러일으켰습니다. 이 아이디어의 핵심은 목표 지향적인 방식으로 시각적 표현을 형성, 유지 및 조작하는 능력입니다. 이 능력을 평가하고 탐구하기 위해, 우리는 프론티어 모델들이 어려워할 만한 난이도로 조정되었으며 시각적 해결이 용이한 절차적이고 계층화된 다단계 추론 문제 세트인 MentisOculi를 개발했습니다. 잠재 토큰에서 명시적으로 생성된 이미지에 이르는 다양한 시각적 전략을 평가한 결과, 이러한 전략들이 일반적으로 성능을 향상시키지 못한다는 사실을 발견했습니다. 특히 UMM에 대한 분석은 치명적인 한계를 드러냅니다. 모델들이 작업을 해결할 수 있는 텍스트 추론 능력을 갖추고 있고 때로는 정확한 시각 자료를 생성할 수도 있지만, 누적되는 생성 오류로 인해 어려움을 겪으며 심지어 정답(ground-truth) 시각화조차 활용하지 못한다는 점입니다. 우리의 연구 결과는 시각적 사고가 가진 내재적 매력에도 불구하고, 아직 모델의 추론에는 도움이 되지 않음을 시사합니다. MentisOculi는 다양한 모델 제품군 전반에서 이러한 격차를 분석하고 해소하기 위한 필수적인 기반을 마련합니다.

Original Abstract

Frontier models are transitioning from multimodal large language models (MLLMs) that merely ingest visual information to unified multimodal models (UMMs) capable of native interleaved generation. This shift has sparked interest in using intermediate visualizations as a reasoning aid, akin to human mental imagery. Central to this idea is the ability to form, maintain, and manipulate visual representations in a goal-oriented manner. To evaluate and probe this capability, we develop MentisOculi, a procedural, stratified suite of multi-step reasoning problems amenable to visual solution, tuned to challenge frontier models. Evaluating visual strategies ranging from latent tokens to explicit generated imagery, we find they generally fail to improve performance. Analysis of UMMs specifically exposes a critical limitation: While they possess the textual reasoning capacity to solve a task and can sometimes generate correct visuals, they suffer from compounding generation errors and fail to leverage even ground-truth visualizations. Our findings suggest that despite their inherent appeal, visual thoughts do not yet benefit model reasoning. MentisOculi establishes the necessary foundation to analyze and close this gap across diverse model families.

2 Citations
0 Influential
3 Altmetric
17.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!