MVP-Nav: 다층 가치 지도 기반 탐색기
MVP-Nav: Multi-layer Value Map Planner Navigator
RGB 이미지만으로 객체 목표 지점까지 이동하는 것은 에이전트에게 근본적인 어려움을 안겨줍니다. 명시적인 깊이 정보가 없으면 심각한 물리적 불확실성과 의미-물리적 일치성 문제가 발생하기 때문입니다. 기존 방법들은 주로 고수준의 의미 기반 추론에 의존하거나, 명시적인 물리적 제약 없이 엔드 투 엔드 정책을 학습하여, 의미적으로는 타당하지만 물리적으로 안전하지 않은 행동으로 이어지는 경우가 많습니다. 본 논문에서는 MVP-Nav라는 물리 인지 RGB 기반 탐색 프레임워크를 제안합니다. MVP-Nav는 3D 기초 모델을 활용하여 단일 카메라 관찰로부터 명시적인 물리적 공간 정보를 재구성하고, 이를 통해 인식, 계획 및 제어를 실제 3D 세계와 일치시킵니다. 특히, 2D 의미 인스턴스를 3D 경계 상자로 투영하여 전역적인 공간 의미 표현을 생성합니다. 또한, 고수준의 의미 기반 추론과 저수준의 물리적 제약을 통합하기 위해, 의미 우선순위와 재구성된 기하학 정보를 공유 비용 공간에 통합하는 다층 가치 지도(MVM)를 도입했습니다. 이를 통해 물리적으로 기반한 기하학적 계획이 가능합니다. 객체 목표 지점 탐색 벤치마크에서 진행된 광범위한 실험 결과, MVP-Nav는 기존의 깊이 정보가 없는 방법들을 크게 능가하며 최첨단 성능을 달성했습니다. 이는 구조화된 물리적 사전 지식이 활성 깊이 센서의 부재를 효과적으로 보완할 수 있음을 입증합니다.
Zero-shot Object Goal Navigation (ZSON) with RGB-only perception poses a fundamental challenge for embodied agents, as the absence of explicit depth information introduces severe physical uncertainty and semantic-physical misalignment. Existing approaches either rely on high-level semantic reasoning without geometric grounding or learn end-to-end policies that lack explicit physical constraints, often resulting in semantically plausible but physically unsafe behaviors. In this paper, we propose MVP-Nav, a physical-aware RGB-only navigation framework that aligns perception, planning, and control with the real 3D world. MVP-Nav reconstructs explicit physical occupancy from monocular observations by leveraging 3D foundation models to project 2D semantic instances into 3D oriented bounding boxes, forming a global spatial semantic representation. To unify high-level semantic reasoning and low-level physical constraints, we introduce a Multi-layer Value Map (MVM) that integrates semantic priorities and reconstructed geometry into a shared cost space, enabling physically grounded geometric planning. Extensive experiments on zero-shot object navigation benchmarks demonstrate that MVP-Nav significantly outperforms existing depth-free methods, achieving state-of-the-art performance and validating that structured physical priors can effectively compensate for the absence of active depth sensors.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.