2607.14514v1 Jul 16, 2026 cs.CV

VTM-Nav: 계층적 시각-위상 메모리를 활용한 에피소드 간 객체-목표 탐색

VTM-Nav: Hierarchical Visual-Topological Memory for Cross-Episode Object-Goal Navigation

Changsheng Xu
Changsheng Xu
Citations: 238
h-index: 7
Yifan Xu
Yifan Xu
Citations: 490
h-index: 6
Xiaoran Xu
Xiaoran Xu
Citations: 23
h-index: 2
Tianyue Xue
Tianyue Xue
Citations: 0
h-index: 0
Xuanran Dong
Xuanran Dong
Citations: 0
h-index: 0
Yupeng Wu
Yupeng Wu
Citations: 1
h-index: 1
Xiaoshan Yang
Xiaoshan Yang
Citations: 24
h-index: 4

객체-목표 탐색은 로봇이 실내 환경에서 특정 객체의 인스턴스를 찾고 도달하는 것을 요구합니다. 최근의 학습 없이 동작하는 방법들은 시각-언어 모델(VLM)을 활용하여 개방형 어휘 기반의 의미적 추론을 수행하지만, 일반적으로 각 에피소드마다 모든 장면 관련 상태를 초기화하는 에피소드 프로토콜 하에서 평가됩니다. 본 연구에서는 Cross-Episode Object-Goal Navigation이라는 새로운 방법을 제시합니다. 이 방법은 로봇이 동일한 장면 내에서 반복적으로 작동하며, 스스로 획득한 경험만을 유지하고 모델 파라미터를 고정합니다. 경험 재사용을 지원하기 위해, 학습 없이 동작하는 VLM 탐색 프레임워크인 exttt{method}를 제안합니다. exttt{method}는 지속적인 계층적 시각-위상 메모리(VTM)를 사용하여 장면 지식을 방과 객체 수준에서 구성하고, 거친 단계부터 세밀한 단계에 이르기까지 매칭을 통해 관련 경험을 검색하며, 현재 관찰 결과와 일치할 때만 메모리를 소프트 가이드로 제공합니다. 또한, 보수적인 실행 제어 장치를 추가하여 진동, 막힌 동작 및 조기 정지를 완화합니다. 본 연구에서는 통제된 동일 장면 프로토콜 하에서 exttt{method}를 HM3D v0.1, HM3D v0.2, 그리고 MP3D의 세 가지 벤치마크 데이터셋에 대해 평가하고, VLM 백본과 액션 파이프라인을 동일하게 유지하면서 강화된 WMNav 기준 모델과 비교했습니다. exttt{method}는 모든 세 가지 벤치마크에서 최상의 성능을 달성했으며, 이는 구조화된 시각-위상 경험 재사용의 효과성과 다양한 데이터셋에서의 견고성을 입증합니다.

Original Abstract

Object-goal navigation requires an embodied agent to locate and reach an instance of a specified object category in an indoor environment. Recent training-free approaches leverage vision-language models (VLMs) for open-vocabulary semantic reasoning, but are typically evaluated under an episodic protocol that resets all scene-specific state after each episode. We introduce Cross-Episode Object-Goal Navigation, in which an agent repeatedly operates in the same scene, retains only self-acquired experience, and keeps its model parameters fixed. To support experience reuse, we present \method, a training-free VLM navigation framework with a persistent hierarchical Visual-Topological Memory (VTM). The VTM organizes scene knowledge at room and object levels and retrieves relevant experience through coarse-to-fine matching, providing memory as soft guidance only when it agrees with current observations. A conservative execution guard further mitigates oscillations, blocked motions, and premature stopping. Under a controlled same-scene protocol, we evaluate \method{} on three benchmarks, HM3D v0.1, HM3D v0.2, and MP3D, and compare it with a strengthened WMNav baseline augmented with cross-episode textual memory, while keeping the VLM backbone and action pipeline identical. \method{} achieves the best performance across all three benchmarks, demonstrating the effectiveness and robustness of structured visual-topological experience reuse across datasets.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!