SpatialText: 대규모 언어 모델의 공간 이해를 위한 순수 텍스트 기반 인지 벤치마크
SpatialText: A Pure-Text Cognitive Benchmark for Spatial Understanding in Large Language Models
진정한 공간 추론은 표면적인 언어 연관성을 처리하는 것 이상으로, 일관성 있는 내부 공간 표현을 구성하고 조작하는 능력에 의존하며, 이는 종종 정신 모델로 개념화됩니다. 대규모 언어 모델은 다양한 분야에서 뛰어난 능력을 보이지만, 기존 벤치마크는 이러한 고유한 공간 인지 능력을 통계적 언어적 휴리스틱과 분리하는 데 실패합니다. 또한, 다중 모드 평가에서는 종종 진정한 공간 추론이 시각적 인식과 혼동됩니다. 모델이 유연한 공간 정신 모델을 구축하는지 체계적으로 조사하기 위해, 우리는 이론에 기반한 진단 프레임워크인 SpatialText를 소개합니다. SpatialText는 단순히 데이터 세트로 기능하는 것이 아니라, 이중 소스 방법론을 통해 텍스트 기반 공간 추론을 분리합니다. SpatialText는 실제 3D 실내 환경에 대한 인간이 주석을 단 자연스러운 모호성, 시점 변화 및 기능적 관계를 포착한 설명과, 형식적인 공간 추론 및 인식적 경계를 탐구하도록 설계된 코드로 생성된 논리적으로 정확한 장면을 통합합니다. 최첨단 모델에 대한 체계적인 평가는 근본적인 표현적 한계를 드러냅니다. 모델은 명시적인 공간 사실을 검색하고 전역적, 이질 좌표계를 사용하여 작동하는 데 능숙함을 보여주지만, 공심적 관점 변환 및 지역 참조 프레임 추론에서 중요한 오류를 보입니다. 이러한 체계적인 오류는 현재 모델이 일관되고 검증 가능한 내부 공간 표현을 구축하는 대신 언어적 공존 휴리스틱에 크게 의존한다는 강력한 증거를 제공합니다. 따라서 SpatialText는 인공 공간 지능의 인지적 경계를 진단하는 엄격한 도구로 사용될 수 있습니다.
Genuine spatial reasoning relies on the capacity to construct and manipulate coherent internal spatial representations, often conceptualized as mental models, rather than merely processing surface linguistic associations. While large language models exhibit advanced capabilities across various domains, existing benchmarks fail to isolate this intrinsic spatial cognition from statistical language heuristics. Furthermore, multimodal evaluations frequently conflate genuine spatial reasoning with visual perception. To systematically investigate whether models construct flexible spatial mental models, we introduce SpatialText, a theory-driven diagnostic framework. Rather than functioning simply as a dataset, SpatialText isolates text-based spatial reasoning through a dual-source methodology. It integrates human-annotated descriptions of real 3D indoor environments, which capture natural ambiguities, perspective shifts, and functional relations, with code-generated, logically precise scenes designed to probe formal spatial deduction and epistemic boundaries. Systematic evaluation across state-of-the-art models reveals fundamental representational limitations. Although models demonstrate proficiency in retrieving explicit spatial facts and operating within global, allocentric coordinate systems, they exhibit critical failures in egocentric perspective transformation and local reference frame reasoning. These systematic errors provide strong evidence that current models rely heavily on linguistic co-occurrence heuristics rather than constructing coherent, verifiable internal spatial representations. SpatialText thus serves as a rigorous instrument for diagnosing the cognitive boundaries of artificial spatial intelligence.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.