AHA-WAM: 관찰 기반 컨텍스트 라우팅을 통한 비동기적 환경 적응형 세계-행동 모델링
AHA-WAM:Asynchronous Horizon-Adaptive World-Action Modeling with Observation-Guided Context Routing
세계-행동 모델은 로봇 조작 분야에서 유망한 패러다임으로 부상했으며, 시각적인 장면의 역학 및 행동을 동시에 모델링하여 정책 학습에 물리적 사전 지식을 주입합니다. 그러나 기존의 세계-행동 모델들은 세계 예측과 행동 실행을 동일한 시간 해상도로 결합하기 때문에, 세계 분기는 불필요하고 정보량이 적은 단기 프레임 변화를 모델링해야 하는 제약을 받습니다. 본 연구에서는 세계 예측과 행동 실행을 동일한 시간 리듬에 묶는 것이 비디오 분기의 잠재력을 충분히 활용하지 못할 수 있다고 가정합니다. 따라서, 본 연구에서는 이와 같은 시간적 비대칭성을 고려하여 설계된, Dual Diffusion Transformer (DiT) 아키텍처를 기반으로 하는 비동기적 환경 적응형 세계-행동 모델인 AHA-WAM을 제안합니다. AHA-WAM은 비디오 DiT를 저주파 세계 플래너로 구현하여 과거 관찰로부터 연속적인 키-값 메모리를 유지하고, 재사용 가능한 레이어별 잠재 컨텍스트 인코딩을 통해 장기적인 장면 진화를 표현합니다. 동시에, 고주파 액션 DiT는 이 컨텍스트를 레이어별 공동 어텐션을 통해 활용하면서 짧은 행동 단위를 폐루프 방식으로 실행합니다. 비동기적 실행을 지원하기 위해, 본 연구에서는 호라이즌 적응 오프셋 학습과 관찰 기반 비디오-컨텍스트 라우팅 (OVCR)을 도입했습니다. 이를 통해 액션 전문가가 장기적인 세계 컨텍스트를 활용하면서도 비디오 DiT를 다시 실행하지 않고 실시간 실행 상태에 대응할 수 있습니다. RoboTwin 및 실제 로봇 조작 작업에서의 실험 결과는 AHA-WAM이 로봇 데이터 사전 훈련 없이 최첨단 성능을 달성함을 보여줍니다. 구체적으로, RoboTwin에서 평균 성공률 92.80%를, 실제 작업 4가지에서 78.3%의 성공률을 기록했으며, Fast-WAM에 비해 4.59배 빠른 24.17 Hz의 폐루프 제어를 달성했습니다.
World-action models have emerged as a promising paradigm for robot manipulation, jointly modeling visual scene dynamics and actions to inject physical priors into policy learning. However, existing world-action models couple world prediction and action execution at the same temporal resolution, forcing the world branch to model near-term frame variations that are redundant and weakly informative. We posit that strictly binding world prediction and action execution to the same temporal rhythm may underutilize the potential of the video branch for embodied control. Therefore, we propose AHA-WAM, an Asynchronous Horizon-Adaptive World-Action Model built on a dual Diffusion Transformer (DiT) architecture that reorganizes world-action modeling around this temporal asymmetry. AHA-WAM instantiates the video DiT as a low-frequency world planner that maintains rolling key-value memory over past observations and exposes reusable layerwise latent context encoding long-horizon scene evolution, while a high-frequency action DiT executes short action chunks in closed loop by querying this context through layerwise joint attention. To support asynchronous execution, we introduce horizon-adaptive offset training and Observation-Guided Video-Context Routing (OVCR), which together let the action expert exploit long-horizon world context while remaining responsive to real-time execution state without rerunning the video DiT. Experiments on RoboTwin and real-world manipulation tasks show that AHA-WAM achieves state-of-the-art performance without any robot-data pretraining, attaining 92.80% average success on RoboTwin and 78.3% success across 4 real-world tasks, while reaching 24.17 Hz closed-loop control with a 4.59x speedup over Fast-WAM.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.