2608.03244v1 Aug 04, 2026 cs.AI

UniNav: 시각적 내비게이션을 위한 통합 세계-행동 확산 모델

UniNav: A Unified World-Action Diffusion Model for Visual Navigation

Changhao Chen
Changhao Chen
Citations: 13
h-index: 2
Yueru Luo
Yueru Luo
Citations: 301
h-index: 7
Changqing Zhou
Changqing Zhou
Citations: 8
h-index: 2
Zeyu Jiang
Zeyu Jiang
Citations: 21
h-index: 3

이미지 기반 목표 지향 시각적 내비게이션은 물리적 에이전트의 기본적인 능력입니다. 기존 내비게이션 정책은 효율적으로 경로 예측을 수행하지만, 시각적인 예측 능력이 부족하며, 내비게이션 세계 모델은 미래 관찰 내용을 예측할 수 있지만, 종종 비용이 많이 드는 계획 실행 과정을 필요로 합니다. 본 논문에서는 UniNav라는 통합된 세계-행동 모델을 제시합니다. UniNav는 단일 확산 과정을 통해 미래 시각적 관찰 내용과 연속적인 경로 지점을 생성합니다. UniNav는 과거 프레임과 목표 이미지 정보를 바탕으로, 하나의 트랜스포머 내에서 시각적 토큰과 경로 토큰을 동시에 제거하여, 미래 예측과 행동 생성 기능을 통합된 프레임워크로 제공합니다. 공간적 위치 정보의 정확도를 높이기 위해, 기하학적인 정보를 고려한 카메라 토큰을 포함했습니다. 또한, UniNav는 경로 정보가 있는 내비게이션 데이터와 비디오 데이터만을 사용한 학습을 통해, 다양한 비디오 데이터를 활용하여 모델 성능을 향상시켰습니다. 이러한 통합된 프레임워크를 기반으로, 우리는 두 가지 변형 모델을 제시합니다: UniNav-Full은 해석 가능한 미래 관찰 내용과 해당 경로 지점을 동시에 예측하며, UniNav-Fast는 추론 과정에서 미래 이미지 토큰을 제거하여 효율적인 경로 예측을 수행합니다. 내비게이션 벤치마크 실험 결과, UniNav는 모든 데이터셋에서 ATE (Average Trajectory Error) 측면에서 가장 강력한 기준 모델보다 우수한 성능을 보였습니다. UniNav-Fast는 단일 단계 추론으로 0.1초의 지연 시간을 가지며, 상당한 정확도 손실 없이 빠른 경로 예측이 가능합니다. 관련 코드는 공개될 예정입니다.

Original Abstract

Image-goal visual navigation is a fundamental capability for embodied agents. Existing navigation policies efficiently predict waypoint trajectories but lack visual foresight, while navigation world models can anticipate future observations but often require costly planning rollouts. We present UniNav, a unified world-action model that generates future visual observations and continuous waypoint trajectories through a single diffusion process. Given history frames and a goal image, UniNav jointly denoises visual and waypoint tokens within a single transformer, unifying future prediction and action generation in a shared framework. To improve spatial grounding, we incorporate geometry-aware camera tokens. We also train on both trajectory-labeled navigation data and video-only data, enabling the model to benefit from diverse videos without waypoint annotations. Based on this unified framework, we introduce two variants: UniNav-Full jointly predicts interpretable future observations and their corresponding trajectories, while UniNav-Fast removes future-image tokens at inference for efficient trajectory prediction. Experiments on navigation benchmarks show that UniNav outperforms the strongest baseline in ATE across all datasets. With one-step inference, UniNav-Fast achieves a latency of 0.1s without a substantial accuracy drop. Code will be released.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!