LatentPilot: 장면 인지 능력을 갖춘 시각-언어 기반 탐색: 잠재 시각적 추론을 통한 미래 예측
LatentPilot: Scene-Aware Vision-and-Language Navigation by Dreaming Ahead with Latent Visual Reasoning
기존의 시각-언어 기반 탐색(VLN) 모델은 주로 과거 및 현재의 시각 정보를 활용하여 추론하지만, 행동으로 인해 발생하는 미래의 시각적 변화는 대부분 고려하지 않습니다. 그 결과, 이러한 모델은 행동과 시각 세계의 변화 간의 인과 관계를 효과적으로 이해하지 못하여, 안정적인 의사 결정을 내리는 데 한계가 있습니다. 반면, 인간은 행동-동역학적 인과 관계를 활용하여 가까운 미래를 상상함으로써 환경 이해 능력과 탐색 선택 능력을 향상시킵니다. 이러한 인간의 능력을 모방하여, 우리는 미래의 시각 정보를 학습 과정에서 귀중한 데이터 소스로 활용하여 행동에 따른 시각적 동역학을 학습하는 새로운 패러다임인 LatentPilot을 제안합니다. 이때, 추론 과정에서는 미래 프레임을 사용할 필요가 없습니다. 구체적으로, 우리는 온-정책 트랙터를 반복적으로 수집하고 모델을 재학습하는 플라이휠 방식의 학습 메커니즘을 제안합니다. 이 메커니즘은 에이전트의 행동 분포에 더 잘 맞도록 설계되었으며, 에이전트가 지나치게 벗어날 경우 전문가의 개입을 통해 보정합니다. LatentPilot은 또한 명시적인 감독 없이 시각적 잠재 토큰을 학습합니다. 이러한 잠재 토큰은 연속적인 잠재 공간에서 전역적으로 상호 작용하며, 각 단계에서 현재의 출력과 다음 입력으로 사용되어 에이전트가 미래를 예측하고 행동이 후속 시각 정보에 미치는 영향을 추론할 수 있도록 합니다. R2R-CE, RxR-CE 및 R2R-PE 벤치마크에서 새로운 최고 성능(SOTA) 결과를 달성했으며, 다양한 환경에서의 실제 로봇 테스트를 통해 LatentPilot이 환경-행동 동역학을 더 잘 이해하는 것을 입증했습니다. 프로젝트 페이지: https://abdd.top/latentpilot/
Existing vision-and-language navigation (VLN) models primarily reason over past and current visual observations, while largely ignoring the future visual dynamics induced by actions. As a result, they often lack an effective understanding of the causal relationship between actions and how the visual world changes, limiting robust decision-making. Humans, in contrast, can imagine the near future by leveraging action-dynamics causality, which improves both environmental understanding and navigation choices. Inspired by this capability, we propose LatentPilot, a new paradigm that exploits future observations during training as a valuable data source to learn action-conditioned visual dynamics, while requiring no access to future frames at inference. Concretely, we propose a flywheel-style training mechanism that iteratively collects on-policy trajectories and retrains the model to better match the agent's behavior distribution, with an expert takeover triggered when the agent deviates excessively. LatentPilot further learns visual latent tokens without explicit supervision; these latent tokens attend globally in a continuous latent space and are carried across steps, serving as both the current output and the next input, thereby enabling the agent to dream ahead and reason about how actions will affect subsequent observations. Experiments on R2R-CE, RxR-CE, and R2R-PE benchmarks achieve new SOTA results, and real-robot tests across diverse environments demonstrate LatentPilot's superior understanding of environment-action dynamics in scene. Project page:https://abdd.top/latentpilot/
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.