상상에 언제 신뢰해야 하는가: 월드 액션 모델을 위한 적응적 행동 실행
When to Trust Imagination: Adaptive Action Execution for World Action Models
최근 월드 액션 모델(WAM)은 미래의 시각적 관찰과 행동을 동시에 예측하여 로봇 조작 분야에서 유망한 패러다임으로 부상했습니다. 그러나 현재의 WAM은 일반적으로 모델 추론 후 고정된 수의 예측된 행동을 실행하며, 이로 인해 로봇은 상상된 미래가 실제 물리적 실행과 일치하는지 여부를 알 수 없습니다. 본 연구에서는 적응적 WAM 실행을 미래-실제 검증 문제로 정의합니다. 즉, WAM이 예측하는 미래가 신뢰할 수 있을 때는 더 긴 시간 동안 행동을 실행하고, 현실이 상상과 달라지는 경우에는 더 빨리 재계획해야 합니다. 이를 위해, 우리는 미래 포워드 다이내믹스 인과적 주의(FFDC)라는 경량 검증기를 제안합니다. FFDC는 예측된 미래 행동, 예측된 시각적 다이내믹, 실제 관찰, 그리고 언어 명령을 함께 고려하여, 남은 행동 실행이 여전히 신뢰할 수 있는지 추정합니다. FFDC는 예측-관찰 일관성을 통해 적응적인 행동 청크 크기를 결정하며, 이를 통해 장기 실행의 효율성을 유지하면서도 접촉이 빈번하거나 어려운 단계에서 반응성을 회복할 수 있습니다. 또한, 우리는 적응적 실행을 위한 장기 궤적 커버리지를 향상시키기 위해 Mixture-of-Horizon Training을 도입했습니다. RoboTwin 벤치마크 및 실제 환경에서의 실험 결과는 제안된 방법이 강력한 안정성-효율성 균형을 제공함을 보여줍니다. RoboTwin에서는 WAM 순방향 패스를 69.10% 줄이고 실행 시간을 34.02% 단축하며 성공률을 2.54% 향상시켰습니다. 실제 환경에서는 성공률을 35% 향상시켰습니다.
World Action Models (WAMs) have recently emerged as a promising paradigm for robotic manipulation by jointly predicting future visual observations and future actions. However, current WAMs typically execute a fixed number of predicted actions after each model inference, leaving the robot blind to whether the imagined future remains consistent with the actual physical rollout. In this work, we formulate adaptive WAM execution as a future-reality verification problem: the robot should execute longer when the WAM-predicted future remains reliable, and replan earlier when reality deviates from imagination. To this end, we propose Future Forward Dynamics Causal Attention (FFDC), a lightweight verifier that jointly reasons over predicted future actions, predicted visual dynamics, real observations, and language instructions to estimate whether the remaining action rollout can still be trusted. FFDC enables adaptive action chunk sizes as an emergent consequence of prediction-observation consistency, preserving the efficiency of long-horizon execution while restoring responsiveness in contact-rich or difficult phases. We further introduce Mixture-of-Horizon Training to improve long-horizon trajectory coverage for adaptive execution. Experiments on the RoboTwin benchmark and in the real world demonstrate that our method achieves a strong robustness-efficiency trade-off: on RoboTwin, it reduces WAM forward passes by 69.10% and execution time by 34.02%, while improving success rate by 2.54% over the short-chunk baseline; in real-world experiments, it improves success rate by 35%.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.