SpaceVLN: 온라인 공간 인지 기억 및 추론을 갖춘 제로샷 비전-언어 내비게이션 에이전트
SpaceVLN: A Zero-Shot Vision-and-Language Navigation Agent with Online Spatial Cognitive Memory and Reasoning
연속적인 환경에서의 비전-언어 내비게이션은 에이전트가 언어 지시를 따르기 위해 이전에 보지 못한 환경의 공간 구조를 이해해야 합니다. 기초 모델은 특정 작업에 대한 정책 훈련 없이 제로샷 내비게이션으로 이어질 수 있는 유망한 방법을 제시했지만, 많은 내비게이션 시스템은 여전히 로컬 시각적 단서와 선형 히스토리 기반 추론에 의존하며, 탐색된 영역, 이동 경로, 랜드마크 및 이들 간의 공간 관계라는 내비게이션의 중요한 측면을 간과합니다. 본 논문에서는 공간 인지 기억과 작업 지향적인 공간 추론을 기반으로 하는 내비게이션 에이전트인 SpaceVLN을 제안합니다. 특히, SpaceVLN은 검증 가능한 공간-랜드마크 단계를 중심으로 계획 및 실행이 구성된 효율적인 단계별 폐루프 프레임워크를 도입합니다. 내비게이션 과정에서 에이전트는 탐색된 영역을 공간 웨이포인트로 점진적으로 추상화하고, 하위 작업에 기반한 랜드마크 증거를 동적으로 유지하여 진행 상황 파악 및 공간 관계 이해를 위한 계층적 공간 인지 기억을 형성합니다. 이러한 기억을 바탕으로 Spatial-CoT는 작업 진행 추론을 공간 인식, 분석 및 예측과 통합하여 임베디드 내비게이션을 위한 작업 지향적인 공간 추론을 가능하게 합니다. 통일된 단계 인터페이스를 통해 SpaceVLN은 특정 작업에 대한 정책 훈련 없이도 비전-언어 내비게이션과 객체-목표 내비게이션 모두를 통합된 제로샷 환경에서 처리할 수 있습니다. R2R-CE, RxR-CE, GN-Bench 및 HM3D-OVON 데이터셋에서 SpaceVLN은 최첨단 제로샷 성능을 달성했으며, 실제 로봇 배포는 그 적용 가능성을 더욱 입증합니다. 이러한 결과는 공간 인지 기억과 작업 지향적인 공간 추론이 보다 강력한 임베디드 내비게이션 에이전트를 구축하기 위한 실용적인 기반임을 강조합니다.
Vision-and-Language Navigation in continuous environments requires agents to understand the spatial structure of previously unseen environments in order to follow language instructions. Although foundation models have opened a promising path toward zero-shot navigation without task-specific policy training, many navigators still rely on local visual cues and linear history-based reasoning, overlooking the spatial nature of navigation across explored regions, traversed paths, landmarks, and their spatial relations. In this paper, we propose SpaceVLN, a navigation agent built around Spatial Cognitive Memory and Task-Guided Spatial Reasoning. Specifically, SpaceVLN introduces an efficient stagewise closed-loop framework where planning and execution are organized around verifiable space--landmark stages. During navigation, the agent progressively abstracts explored regions into Spatial Waypoints and dynamically maintains subtask-grounded landmark evidence, forming a hierarchical Spatial Cognitive Memory for progress localization and spatial-relation understanding. Built on this memory, Spatial-CoT integrates task-progress reasoning with spatial perception, analysis, and prediction, enabling Task-Guided Spatial Reasoning for embodied navigation. The unified stage interface enables SpaceVLN to address both Vision-and-Language Navigation and Object-Goal Navigation under a unified zero-shot setting, without task-specific policy training. Across R2R-CE, RxR-CE, GN-Bench, and HM3D-OVON, SpaceVLN achieves state-of-the-art zero-shot performance, and real-robot deployment further validates its applicability. These results highlight Spatial Cognitive Memory and Task-Guided Spatial Reasoning as a practical foundation for stronger embodied navigation agents.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.