보이지 않고 예측하기: 월드 액션 모델을 위한 잠재적 미래
Foresight Without Seeing: Latent Futures for World Action Models
월드 액션 모델(WAM)은 로봇의 행동 생성과 함께 미래 시각적 예측을 결합하여, 상호작용 동안 물리 세계가 어떻게 변화하는지 모델링할 수 있도록 합니다. 기존 WAM들은 예측 동역학이 행동 경로에 어떻게 노출되는지에 따라 다릅니다. 명시적인 미래를 사용하는 WAM은 예측된 장면의 변화에 직접 접근할 수 있지만, 반복적인 비디오 디노이징으로 인해 상당한 추론 비용이 발생합니다. 반대로, 직접 정책을 사용하는 WAM은 현재 관측으로부터 효율적으로 행동을 예측하지만, 액션 DiT에 예측 동역학을 노출하기 위한 명시적인 런타임 인터페이스가 부족합니다. 이러한 격차를 해소하기 위해, 우리는 ForeWAM이라는 동역학 기반의 직접 정책 WAM을 제안합니다. ForeWAM은 미래 비디오 디코딩 없이 행동 생성에 필요한 예측 컨텍스트를 제공합니다. 핵심 구성 요소인 Future-KV는 현재 시각적 잠재 정보와 확률적인 미래 슬롯에 대해 단일 Video DiT 프리필 과정을 수행하고, 그 결과로 얻어진 레이어별 키-값 상태를 행동 디노이징 과정 전체에서 재사용합니다. 또한, 동결된 잠재 액션 티처의 감독 하에 생성되는 동역학 레지스터를 도입하여, 암시적인 미래 상태가 상호작용으로 인한 변화(예: 물체의 움직임, 접촉 변화 및 작업 진행)를 포착하도록 유도합니다. 실제 미래 관측 정보와 티저는 학습 과정에서만 사용되며, 배포 시에는 필요하지 않으며 미래 비디오 생성이 전혀 이루어지지 않습니다. 로봇의 실제 데이터를 활용한 사전 훈련 없이, ForeWAM의 표준 버전과 가속화된 버전은 각각 LIBERO 데이터셋에서 평균 성공률 96.7%와 96.9%를 달성합니다. 또한, 표준 버전은 LIBERO-Plus 데이터셋에서도 61.6%의 성공률을 보입니다. 이러한 결과는 직접 정책 WAM이 효율적인 행동 예측을 유지하면서 미래 관측 정보를 명시적으로 생성하지 않고도 예측 동역학을 행동 경로에 노출할 수 있음을 보여줍니다.
World Action Models (WAMs) couple future visual prediction with robot action generation, enabling policies to model how the physical world evolves during interaction. Existing WAMs differ in how predictive dynamics are exposed to the action pathway. Explicit-future WAMs provide direct access to predicted scene evolution, but incur substantial inference costs from iterative video denoising. In contrast, direct-policy WAMs efficiently predict actions from the current observation but lack an explicit inference-time interface for exposing predictive dynamics to the Action DiT. To bridge this gap, we propose ForeWAM, a dynamics-conditioned direct-policy WAM that provides predictive context for action generation without decoding future videos. At its core, Future-KV performs a single Video DiT prefill over the current visual latent and stochastic future slots, and reuses the resulting layer-wise key-value states throughout action denoising. We further introduce dynamics registers supervised by a frozen latent action teacher, encouraging the implicit future states to capture interaction-induced transitions such as object motion, contact changes, and task progress. Ground-truth future observations and the teacher are used only during training; deployment requires neither and performs no future video generation. Without embodied robot data pretraining, the standard and accelerated variants of ForeWAM achieve average success rates of 96.7% and 96.9% on LIBERO, respectively. The standard variant further achieves 61.6% success on LIBERO-Plus. These results demonstrate that direct-policy WAMs can retain efficient action prediction while exposing predictive dynamics to the action pathway without explicitly generating future observations.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.