2606.09669v1 Jun 08, 2026 cs.AI

SpatialWorld: 실제 환경 작업에서 다중 모드 에이전트의 상호작용적 공간 추론 성능 평가

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks

Hongyi Yuan
Hongyi Yuan
Citations: 6,219
h-index: 19
Zihao Huang
Zihao Huang
Citations: 1,565
h-index: 6
Nan Duan
Nan Duan
Citations: 43
h-index: 3
Bohan Zeng
Bohan Zeng
Citations: 234
h-index: 9
Wenjie Li
Wenjie Li
Citations: 23
h-index: 3
Wentao Zhang
Wentao Zhang
Citations: 38
h-index: 3
Bo Wang
Bo Wang
Citations: 35
h-index: 2
Haoyang Huang
Haoyang Huang
Citations: 397
h-index: 5
Hongcheng Gao
Hongcheng Gao
Citations: 83
h-index: 3
Hailong Qu
Hailong Qu
Citations: 31
h-index: 2
Jingyi Tang
Jingyi Tang
Citations: 28
h-index: 2
Jiahao Wang
Jiahao Wang
Citations: 19
h-index: 2
Hengkang Qiao
Hengkang Qiao
Citations: 0
h-index: 0
Shihong Huang
Shihong Huang
Citations: 24
h-index: 1
Junming Yang
Junming Yang
Citations: 5
h-index: 2
Yi Li
Yi Li
Citations: 78
h-index: 5
Wenbo Li
Wenbo Li
Citations: 259
h-index: 2
Jianhui Liu
Jianhui Liu
Citations: 89
h-index: 2
Oliver Huang
Oliver Huang
Citations: 48
h-index: 3
Guo-Ting Huang
Guo-Ting Huang
Citations: 61
h-index: 1
Yinpeng Dong
Yinpeng Dong
Citations: 218
h-index: 5

공간 추론은 다중 모드 대규모 언어 모델(MLLM)이 물리적인 세계를 인식하고 상호 작용하는 데 필수적인 능력입니다. 그러나 기존의 벤치마크는 주로 수동적인 평가(예: 정적 질의응답) 또는 시뮬레이터에 특화된 파이프라인에 의존하여, 일반적인 상호작용적 공간 이해 능력을 제대로 평가하지 못합니다. 본 논문에서는 복잡한 실제 환경 작업에서 다중 모드 에이전트의 상호작용적 공간 이해 능력을 평가하기 위해 특별히 설계된 통합 벤치마크인 SpatialWorld를 소개합니다. SpatialWorld는 다양한 도메인의 760개의 인간이 직접 작성한 작업을 포함하며, 각 작업은 공유되고 시뮬레이터에 독립적인 프로토콜을 통해 연결된 8가지 이기종 시뮬레이션 백엔드를 사용합니다(예: 일상생활, 여행, 사회적 협력). 에이전트는 제한된 시각 정보만을 활용하여 능동적으로 시각 정보를 수집하고, MLLM에서 사용하는 통일된 텍스트 기반 행동 인터페이스를 통해 의사 결정을 수행해야 합니다. 신뢰성 있는 평가를 위해 각 작업은 인간 검증을 거친 초기 상태, 참조 경로 및 종료 상태 확인 기능을 포함합니다. 15개의 최첨단 에이전트를 평가한 결과, 견고한 공간 추론 능력은 여전히 어려운 과제임이 드러났습니다. 가장 뛰어난 모델인 GPT-5의 평균 작업 성공률(TSR)은 17.4%에 불과했으며, 선도적인 오픈 소스 모델인 Qwen-3.5는 14.1%를 기록했습니다. 추가 분석 결과, 작업 성공 여부와 실행 효율성 간에 명확한 불일치가 있으며, 도메인별 성능 변동이 상당하다는 사실이 밝혀졌습니다. 이러한 능동적 탐색 및 장기 계획 수립의 한계점은 SpatialWorld를 미래의 공간 추론 에이전트를 위한 엄격한 테스트 환경으로 자리매김하게 합니다.

Original Abstract

Spatial reasoning is a foundational capability for multimodal large language models (MLLMs) to perceive and operate within the physical world. However, existing benchmarks predominantly rely on passive evaluation (e.g., static VQA) or simulator-specific pipelines, failing to assess general interactive spatial understanding. We introduce SpatialWorld, a unified benchmark designed specifically for evaluating the interactive spatial understanding of multimodal agents in complex real-world tasks. Integrating eight heterogeneous simulation backends under a shared, simulator-agnostic protocol, SpatialWorld features 760 human-annotated tasks across diverse domains (e.g., household routines, travel, social collaboration). Agents must solve tasks under vision-only partial observability, actively gathering egocentric visual evidence and expressing decisions via a unified, text-based action interface native to MLLMs. For reliable evaluation, each task includes a human-validated initial state, a reference trajectory, and a terminal-state verifier. Evaluating 15 advanced agents reveals that robust spatial task solving remains challenging: the strongest model, GPT-5, achieves an average task success rate (TSR) of only 17.4%, while the leading open-source model, Qwen-3.5, reaches 14.1%. Further analysis exposes a clear mismatch between task success and execution efficiency, alongside substantial domain-specific performance variations. These bottlenecks in active exploration and long-horizon planning position SpatialWorld as a rigorous testbed for future spatial agents.

0 Citations
0 Influential
9.5 Altmetric
47.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!