2606.08992v1 Jun 08, 2026 cs.RO

SpaceVLN: 온라인 공간 인지 기억 및 추론을 갖춘 제로샷 비전-언어 내비게이션 에이전트

SpaceVLN: A Zero-Shot Vision-and-Language Navigation Agent with Online Spatial Cognitive Memory and Reasoning

Chenjia Bai
Chenjia Bai
Citations: 305
h-index: 9
Xuelong Li
Xuelong Li
Citations: 350
h-index: 10
Yu Deng
Yu Deng
Citations: 2
h-index: 1
Pingrui Lai
Pingrui Lai
Citations: 12
h-index: 2
Xinhai Li
Xinhai Li
Citations: 115
h-index: 3
Xiaoheng Deng
Xiaoheng Deng
Citations: 26
h-index: 2
Cheng Sun
Cheng Sun
Citations: 1
h-index: 1
Hua Yang
Hua Yang
Citations: 141
h-index: 4

연속적인 환경에서의 비전-언어 내비게이션은 에이전트가 언어 지시를 따르기 위해 이전에 보지 못한 환경의 공간 구조를 이해해야 합니다. 기초 모델은 특정 작업에 대한 정책 훈련 없이 제로샷 내비게이션으로 이어질 수 있는 유망한 방법을 제시했지만, 많은 내비게이션 시스템은 여전히 로컬 시각적 단서와 선형 히스토리 기반 추론에 의존하며, 탐색된 영역, 이동 경로, 랜드마크 및 이들 간의 공간 관계라는 내비게이션의 중요한 측면을 간과합니다. 본 논문에서는 공간 인지 기억과 작업 지향적인 공간 추론을 기반으로 하는 내비게이션 에이전트인 SpaceVLN을 제안합니다. 특히, SpaceVLN은 검증 가능한 공간-랜드마크 단계를 중심으로 계획 및 실행이 구성된 효율적인 단계별 폐루프 프레임워크를 도입합니다. 내비게이션 과정에서 에이전트는 탐색된 영역을 공간 웨이포인트로 점진적으로 추상화하고, 하위 작업에 기반한 랜드마크 증거를 동적으로 유지하여 진행 상황 파악 및 공간 관계 이해를 위한 계층적 공간 인지 기억을 형성합니다. 이러한 기억을 바탕으로 Spatial-CoT는 작업 진행 추론을 공간 인식, 분석 및 예측과 통합하여 임베디드 내비게이션을 위한 작업 지향적인 공간 추론을 가능하게 합니다. 통일된 단계 인터페이스를 통해 SpaceVLN은 특정 작업에 대한 정책 훈련 없이도 비전-언어 내비게이션과 객체-목표 내비게이션 모두를 통합된 제로샷 환경에서 처리할 수 있습니다. R2R-CE, RxR-CE, GN-Bench 및 HM3D-OVON 데이터셋에서 SpaceVLN은 최첨단 제로샷 성능을 달성했으며, 실제 로봇 배포는 그 적용 가능성을 더욱 입증합니다. 이러한 결과는 공간 인지 기억과 작업 지향적인 공간 추론이 보다 강력한 임베디드 내비게이션 에이전트를 구축하기 위한 실용적인 기반임을 강조합니다.

Original Abstract

Vision-and-Language Navigation in continuous environments requires agents to understand the spatial structure of previously unseen environments in order to follow language instructions. Although foundation models have opened a promising path toward zero-shot navigation without task-specific policy training, many navigators still rely on local visual cues and linear history-based reasoning, overlooking the spatial nature of navigation across explored regions, traversed paths, landmarks, and their spatial relations. In this paper, we propose SpaceVLN, a navigation agent built around Spatial Cognitive Memory and Task-Guided Spatial Reasoning. Specifically, SpaceVLN introduces an efficient stagewise closed-loop framework where planning and execution are organized around verifiable space--landmark stages. During navigation, the agent progressively abstracts explored regions into Spatial Waypoints and dynamically maintains subtask-grounded landmark evidence, forming a hierarchical Spatial Cognitive Memory for progress localization and spatial-relation understanding. Built on this memory, Spatial-CoT integrates task-progress reasoning with spatial perception, analysis, and prediction, enabling Task-Guided Spatial Reasoning for embodied navigation. The unified stage interface enables SpaceVLN to address both Vision-and-Language Navigation and Object-Goal Navigation under a unified zero-shot setting, without task-specific policy training. Across R2R-CE, RxR-CE, GN-Bench, and HM3D-OVON, SpaceVLN achieves state-of-the-art zero-shot performance, and real-robot deployment further validates its applicability. These results highlight Spatial Cognitive Memory and Task-Guided Spatial Reasoning as a practical foundation for stronger embodied navigation agents.

0 Citations
0 Influential
5 Altmetric
25.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!