SpatialWorld: 실제 환경 작업에서 다중 모드 에이전트의 상호작용적 공간 추론 성능 평가
SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks
공간 추론은 다중 모드 대규모 언어 모델(MLLM)이 물리적인 세계를 인식하고 상호 작용하는 데 필수적인 능력입니다. 그러나 기존의 벤치마크는 주로 수동적인 평가(예: 정적 질의응답) 또는 시뮬레이터에 특화된 파이프라인에 의존하여, 일반적인 상호작용적 공간 이해 능력을 제대로 평가하지 못합니다. 본 논문에서는 복잡한 실제 환경 작업에서 다중 모드 에이전트의 상호작용적 공간 이해 능력을 평가하기 위해 특별히 설계된 통합 벤치마크인 SpatialWorld를 소개합니다. SpatialWorld는 다양한 도메인의 760개의 인간이 직접 작성한 작업을 포함하며, 각 작업은 공유되고 시뮬레이터에 독립적인 프로토콜을 통해 연결된 8가지 이기종 시뮬레이션 백엔드를 사용합니다(예: 일상생활, 여행, 사회적 협력). 에이전트는 제한된 시각 정보만을 활용하여 능동적으로 시각 정보를 수집하고, MLLM에서 사용하는 통일된 텍스트 기반 행동 인터페이스를 통해 의사 결정을 수행해야 합니다. 신뢰성 있는 평가를 위해 각 작업은 인간 검증을 거친 초기 상태, 참조 경로 및 종료 상태 확인 기능을 포함합니다. 15개의 최첨단 에이전트를 평가한 결과, 견고한 공간 추론 능력은 여전히 어려운 과제임이 드러났습니다. 가장 뛰어난 모델인 GPT-5의 평균 작업 성공률(TSR)은 17.4%에 불과했으며, 선도적인 오픈 소스 모델인 Qwen-3.5는 14.1%를 기록했습니다. 추가 분석 결과, 작업 성공 여부와 실행 효율성 간에 명확한 불일치가 있으며, 도메인별 성능 변동이 상당하다는 사실이 밝혀졌습니다. 이러한 능동적 탐색 및 장기 계획 수립의 한계점은 SpatialWorld를 미래의 공간 추론 에이전트를 위한 엄격한 테스트 환경으로 자리매김하게 합니다.
Spatial reasoning is a foundational capability for multimodal large language models (MLLMs) to perceive and operate within the physical world. However, existing benchmarks predominantly rely on passive evaluation (e.g., static VQA) or simulator-specific pipelines, failing to assess general interactive spatial understanding. We introduce SpatialWorld, a unified benchmark designed specifically for evaluating the interactive spatial understanding of multimodal agents in complex real-world tasks. Integrating eight heterogeneous simulation backends under a shared, simulator-agnostic protocol, SpatialWorld features 760 human-annotated tasks across diverse domains (e.g., household routines, travel, social collaboration). Agents must solve tasks under vision-only partial observability, actively gathering egocentric visual evidence and expressing decisions via a unified, text-based action interface native to MLLMs. For reliable evaluation, each task includes a human-validated initial state, a reference trajectory, and a terminal-state verifier. Evaluating 15 advanced agents reveals that robust spatial task solving remains challenging: the strongest model, GPT-5, achieves an average task success rate (TSR) of only 17.4%, while the leading open-source model, Qwen-3.5, reaches 14.1%. Further analysis exposes a clear mismatch between task success and execution efficiency, alongside substantial domain-specific performance variations. These bottlenecks in active exploration and long-horizon planning position SpatialWorld as a rigorous testbed for future spatial agents.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.