2606.29908v1 Jun 29, 2026 cs.RO

탐색의 길: 몸체 인식 내비게이션을 위한 공간 인지 세계 행동 모델

Pondering the Way: Spatial-perceiving World Action Model for Embodied Navigation

Guang Chen
Guang Chen
Citations: 282
h-index: 8
Haiyang Sun
Haiyang Sun
Citations: 271
h-index: 7
Bing Wang
Bing Wang
Citations: 264
h-index: 8
Hangjun Ye
Hangjun Ye
Citations: 282
h-index: 8
Zehan Zhang
Zehan Zhang
Citations: 23
h-index: 2
Fang Li
Fang Li
Citations: 186
h-index: 4
Hong Chen
Hong Chen
Citations: 25
h-index: 2
Haiguang Wang
Haiguang Wang
Citations: 292
h-index: 4
Tianhao Lu
Tianhao Lu
Citations: 9
h-index: 2
Hong-Bin Xie
Hong-Bin Xie
Citations: 14
h-index: 2
Daqi Liu
Daqi Liu
Citations: 382
h-index: 11
Longfei Yan
Longfei Yan
Citations: 108
h-index: 2
Yihua Tan
Yihua Tan
Citations: 21
h-index: 1

시각적 내비게이션을 위한 기존의 세계 모델 기반 계획 시스템들은 일반적으로 검증 중심적인 패러다임을 따르며, 목표 의도를 경로 생성과 분리합니다. 이러한 방식은 후보 종속성, 과도한 계산 비용, 그리고 샘플링된 행동과 예측된 시각 정보 간의 불일치 문제를 야기합니다. 이러한 문제점을 해결하기 위해, 우리는 공간 인지 세계 행동 모델 (SWAM)을 제안합니다. SWAM은 목표 중심적인 통합 관찰-행동 생성 프레임워크로, 시작 및 목표 RGB 이미지를 입력으로 받아 단일 단계 추론을 통해 중간 RGB-D 시퀀스와 해당 행동 경로를 동시에 생성하여, 목표 일관성 있는 경로 생성을 촉진하고 공간적 타당성을 향상시킵니다. SWAM은 훈련 과정에서 깊이 가짜 레이블을 활용하여 공간적 사전 지식을 내재화하지만, 추론 시에는 단안 RGB 입력만 필요합니다. 또한, 우리는 동작과 시각 정보 간의 미세한 정렬을 강화하고 다양한 거리에서의 예측 안정성을 확보하기 위해 시각 기반 행동 개선 모듈과 경로 규모 정규화 손실 함수를 추가적으로 도입했습니다. 광범위한 실험 결과는 SWAM이 최첨단 2단계 계획 시스템보다 성공률, 경로 정확도 및 추론 효율성 측면에서 상당한 성능 향상을 보이며, 알려지지 않은 환경에서도 강력한 제로샷 일반화 능력을 보여준다는 것을 입증합니다.

Original Abstract

Existing world model-based planners for visual navigation typically follow a verification-centric paradigm, decoupling goal intent from trajectory synthesis. This approach suffers from candidate dependence, heavy computational overhead, and inconsistencies between sampled actions and predicted visuals. To address these issues, we propose SWAM (Spatial-perceiving World Action Model), a task-centric joint observation-action generation framework. Given start and goal RGB observations, SWAM performs single-pass inference to simultaneously generate intermediate RGB-D sequences and corresponding action trajectories, promoting goal-consistent trajectory generation and improved spatial feasibility. While SWAM leverages depth pseudo-labels during training to internalize spatial priors, it requires only monocular RGB input at inference time. We further introduce a visual-guided action refinement module and a trajectory-scale regularization loss to enforce fine-grained alignment between motion and visual cues while stabilizing predictions across varying distances. Extensive experiments show that SWAM significantly outperforms state-of-the-art two-stage planners in success rate, trajectory accuracy, and inference efficiency, while demonstrating robust zero-shot generalization to unseen environments.

0 Citations
0 Influential
5.5 Altmetric
27.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!