탐색의 길: 몸체 인식 내비게이션을 위한 공간 인지 세계 행동 모델
Pondering the Way: Spatial-perceiving World Action Model for Embodied Navigation
시각적 내비게이션을 위한 기존의 세계 모델 기반 계획 시스템들은 일반적으로 검증 중심적인 패러다임을 따르며, 목표 의도를 경로 생성과 분리합니다. 이러한 방식은 후보 종속성, 과도한 계산 비용, 그리고 샘플링된 행동과 예측된 시각 정보 간의 불일치 문제를 야기합니다. 이러한 문제점을 해결하기 위해, 우리는 공간 인지 세계 행동 모델 (SWAM)을 제안합니다. SWAM은 목표 중심적인 통합 관찰-행동 생성 프레임워크로, 시작 및 목표 RGB 이미지를 입력으로 받아 단일 단계 추론을 통해 중간 RGB-D 시퀀스와 해당 행동 경로를 동시에 생성하여, 목표 일관성 있는 경로 생성을 촉진하고 공간적 타당성을 향상시킵니다. SWAM은 훈련 과정에서 깊이 가짜 레이블을 활용하여 공간적 사전 지식을 내재화하지만, 추론 시에는 단안 RGB 입력만 필요합니다. 또한, 우리는 동작과 시각 정보 간의 미세한 정렬을 강화하고 다양한 거리에서의 예측 안정성을 확보하기 위해 시각 기반 행동 개선 모듈과 경로 규모 정규화 손실 함수를 추가적으로 도입했습니다. 광범위한 실험 결과는 SWAM이 최첨단 2단계 계획 시스템보다 성공률, 경로 정확도 및 추론 효율성 측면에서 상당한 성능 향상을 보이며, 알려지지 않은 환경에서도 강력한 제로샷 일반화 능력을 보여준다는 것을 입증합니다.
Existing world model-based planners for visual navigation typically follow a verification-centric paradigm, decoupling goal intent from trajectory synthesis. This approach suffers from candidate dependence, heavy computational overhead, and inconsistencies between sampled actions and predicted visuals. To address these issues, we propose SWAM (Spatial-perceiving World Action Model), a task-centric joint observation-action generation framework. Given start and goal RGB observations, SWAM performs single-pass inference to simultaneously generate intermediate RGB-D sequences and corresponding action trajectories, promoting goal-consistent trajectory generation and improved spatial feasibility. While SWAM leverages depth pseudo-labels during training to internalize spatial priors, it requires only monocular RGB input at inference time. We further introduce a visual-guided action refinement module and a trajectory-scale regularization loss to enforce fine-grained alignment between motion and visual cues while stabilizing predictions across varying distances. Extensive experiments show that SWAM significantly outperforms state-of-the-art two-stage planners in success rate, trajectory accuracy, and inference efficiency, while demonstrating robust zero-shot generalization to unseen environments.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.