2605.28277v1 May 27, 2026 cs.AI

LLM은 텍스트로부터 세계 모델을 구축하는가? 다국어 공간 추론 진단 연구

Do LLMs Build World Models From Text? A Multilingual Diagnostic of Spatial Reasoning

Xi Xiao
Xi Xiao
Citations: 34
h-index: 1
Zhangquan Chen
Zhangquan Chen
Citations: 213
h-index: 8
Chunlei Meng
Chunlei Meng
Citations: 9
h-index: 2
Zhikai Pan
Zhikai Pan
Citations: 0
h-index: 0
Chih-Ting Liao
Chih-Ting Liao
Citations: 14
h-index: 2
Yitong Qiao
Yitong Qiao
Citations: 25
h-index: 3
Chunrui Liu
Chunrui Liu
Citations: 121
h-index: 5
Xinzhuo Cao
Xinzhuo Cao
Citations: 2
h-index: 1

대규모 언어 모델(LLM)이 순수한 텍스트 설명을 통해 내부적인 공간 세계 모델을 구축하는지에 대한 논쟁은 여전히 진행 중이며, 이러한 기능이 여러 언어로 얼마나 잘 전달되는지는 체계적으로 연구되지 않았습니다. 본 연구에서는 MentalMap이라는 다국어 진단 벤치마크를 소개합니다. MentalMap은 원자적 공간 사실부터 생성적인 세계 그래프 구축까지 여섯 단계의 능력 계층 구조(L0-L5)로 구성되어 있으며, 참조점, 읽기 방향 편향, 추론 노력 배분, 환각 현상에 대한 네 가지 진단 축을 포함합니다. MentalMap은 100개의 ProcTHOR 가정 환경 장면으로 구축되었으며, 8개의 언어학적으로 다양한 언어와 구조화된 텍스트 제어를 포함하고 있으며, 총 1,950개의 평가 항목과 39개의 작업 유형으로 구성되어 있습니다. 다양한 크기와 모델 아키텍처를 가진 13개의 LLM을 평가한 결과, 보편적인 L3 추론의 한계점을 발견했습니다. 즉, 기본 원자적 정확도가 40%를 초과하면 어떤 모델도 시점 기반 추론에서 초기 성능의 절반 이상을 유지하지 못합니다. 이러한 한계는 언어, 크기 및 프롬프트 전략에 관계없이 지속되며, 구조화된 출력 실패와 추론 패턴은 모델마다 크게 다릅니다. 동일한 순수 텍스트 프로토콜 하에서 인간 평가를 수행한 결과, 동일한 실패 패턴이 재현되었으며, 이는 현재 LLM 아키텍처가 아닌 텍스트만으로 작동하는 메모리 제약으로 인해 발생하는 문제임을 시사합니다. 본 연구는 순수 텍스트 기반 공간 추론을 다각적인 세계 모델링 문제로 재정의하고, 향후 다중 모드 및 보조 추론 방법을 활용한 연구를 촉구합니다.

Original Abstract

Whether large language models (LLMs) construct internal spatial world models from pure-text descriptions remains contested, and whether such capabilities transfer across languages has not been systematically studied. We introduce MentalMap, a multilingual diagnostic benchmark with a six-level capability hierarchy (L0-L5) spanning atomic spatial facts to generative world-graph construction, together with four diagnostic axes probing frame of reference, reading-direction bias, reasoning-effort allocation, and hallucination. MentalMap is built from 100 ProcTHOR household scenes, covers eight typologically diverse languages plus a structured-text control, and contains 39 task families across 1,950 evaluation cells. Evaluating thirteen LLMs across scales and model families, we identify a universal L3 reasoning cliff: no model retains even half of its L0 performance on viewpoint reasoning once baseline atomic accuracy exceeds 40%. The cliff persists across languages, scales, and prompting strategies, while structured-output failures and reasoning patterns vary substantially across models. Human evaluation under the identical pure-text protocol reproduces the same failure pattern, suggesting that the bottleneck arises from text-only working memory constraints rather than being specific to current LLM architectures. Our findings reframe pure-text spatial reasoning as a multi-axis world-modeling problem and motivate multimodal and scratchpad-augmented reasoning as future directions.

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!