2606.08952v1 Jun 08, 2026 cs.AI

AlloSpatial: 기반 모델의 공간 추론을 위한 에이전트 기반 프레임워크

AlloSpatial: Agentic Harness Framework for Spatial Reasoning in Foundation Models

Zhenyu Wu
Zhenyu Wu
Citations: 1,521
h-index: 9
Shouwei Ruan
Shouwei Ruan
Citations: 335
h-index: 8
Qihui Zhu
Qihui Zhu
Citations: 12
h-index: 1
Yubin Wang
Yubin Wang
Citations: 4
h-index: 1
Bin Wang
Bin Wang
Citations: 34
h-index: 3
Yuxiang Zhang
Yuxiang Zhang
Citations: 8
h-index: 2
Xingxing Wei
Xingxing Wei
Citations: 271
h-index: 8
Jingzhi Li
Jingzhi Li
Citations: 40
h-index: 2

다중 모드 기반 모델(MFM)은 상당한 발전을 이루었지만, 여전히 실제 세계에 대한 공간 추론 능력에서 취약점을 보입니다. 주요 병목 현상은 로컬 자가 중심 관찰을 전역적이고 객관적인 공간 표현으로 변환하는 데 어려움이 있기 때문입니다. 이를 해결하기 위해, 우리는 기반 모델의 객관적인 공간 인지를 위한 에이전트 기반 프레임워크인 AlloSpatial을 제안합니다. AlloSpatial은 플러그 앤 플레이 방식으로 작동하는 인지 지도 생성 환경인 World2Mind을 도입하여, 자가 중심 관찰을 구조화된 객관적인 사전 정보로 변환합니다. 여기에는 객체 토폴로지, 기하학적 관계, 통행 가능성 및 경로를 쿼리할 수 있는 객관적 공간 트리와 경로 지도가 포함됩니다. AlloSpatial은 또한 노이즈가 많은 재구성 및 불분명한 시각적 증거 하에서 이러한 사전 정보를 신뢰성 있게 활용하기 위해 도구 사용 판단, 모달리티 분리된 단서 수집 및 기하학-의미론 간 조화를 위한 공간 추론 시스템을 도입합니다. 또한, AlloSpatial은 cold-start 강화 학습을 통해 trajectory 수준의 보상을 제공하는 harness-gated 방식으로 Qwen3-VL 모델에 이러한 과정을 내재화했습니다. VSI-Bench와 MindCube 데이터셋에서의 실험 결과, AlloSpatial은 훈련 없이 독점 모델의 성능을 5%에서 18% 향상시켰으며, 객관적 공간 트리(AST)만으로도 시각적 입력이 제거된 경우에도 강력한 공간 추론 능력을 제공합니다. 학습된 AlloSpatial 에이전트는 더 크고 범용적인 모델 및 기존 공간 추론 기술보다 뛰어난 성능을 보여주며, 이는 구조화된 객관적인 표현, 능동적인 도구 사용 및 검증 가능한 추론이 공간적 능력을 갖춘 기반 모델 개발에 유망한 경로임을 시사합니다.

Original Abstract

Multimodal Foundation Models (MFMs) have made substantial progress, yet remain fragile in spatial reasoning over the physical world. A key bottleneck lies in their inability to transform local egocentric observations into a global allocentric spatial representation. To address this, we propose AlloSpatial, an agentic framework for allocentric spatial cognition in foundation models. AlloSpatial introduces World2Mind, a plug-and-play cognitive mapping sandbox that converts egocentric observations into structured allocentric priors, including Allocentric-Spatial Trees and route maps that support querying object topology, geometric relations, passability, and trajectories. To utilize these priors reliably under noisy reconstruction and ambiguous visual evidence, AlloSpatial introduces a Spatial Reasoning Harness for tool-use judgment, modality-decoupled cue collection, and geometry-semantic arbitration. We further internalize this process in Qwen3-VL through cold-start reinforcement learning with a harness-gated trajectory-level reward. Experiments on VSI-Bench and MindCube show that AlloSpatial improves proprietary models by 5%-18% in a training-free setting, while ASTs alone support strong spatial reasoning even when visual inputs are removed. The trained AlloSpatial agents further outperform larger general-purpose models and competitive spatial baselines, suggesting that structured allocentric representations, active tool use, and verifiable reasoning offer a promising route toward spatially capable foundation models.

1 Citations
0 Influential
4.5 Altmetric
23.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!