통합 세계 모델링, 언어 추론 및 행동 합성 위한 월드-언어-액션 모델
World-Language-Action Model for Unified World Modeling, Language Reasoning, and Action Synthesis
본 논문에서는 새로운 유형의 임베디드 기초 모델인 월드-언어-액션 (WLA) 모델을 제안합니다. WLA는 텍스트 지침, 이미지, 로봇 상태를 입력으로 받아 텍스트 하위 작업, 목표 이미지, 로봇 액션을 동시에 예측하며, 월드-액션 모델 (WAM)과 같이 광범위한 1인칭 비디오 데이터를 통해 학습하는 extit{세계 모델링 인터페이스}와 비전-언어-액션 (VLA) 모델처럼 복잡하고 장기적인 작업을 해결하기 위한 extit{언어 추론 능력}을 결합합니다. WLA의 핵심은 WAM에서 사용되는 양방향 디퓨전 트랜스포머 대신, extit{자기 회귀 (AR)} 트랜스포머 백본을 사용하여 extit{다음 상태}를 예측하는데, 이는 extit{의미 수준의} 텍스트 의도와 보완적인 extit{세밀한} 물리적 동역학으로 구성됩니다. 물리적 동역학은 전용 월드 전문가를 기반으로 하는 세계 모델링 목표에 의해 지도 학습되며, 액션 전문가가 상태-액션 상관 관계를 쉽게 파악할 수 있도록 활용됩니다. WLA는 메타 쿼리를 사용하여 세계 예측이 extit{암묵적으로} 행동 생성에 영향을 미치도록 하며, 이를 통해 추론 시 세계 예측 기능을 비활성화할 수 있습니다. 또한, 세계 예측을 활성화하여 테스트 시간에 확장성을 확보하고 로봇 제어를 향상시킬 수도 있습니다. 20억 개의 활성 파라미터를 가진 WLA-0 프로토타입은 NVIDIA RTX 5090에서 추론당 40ms의 성능을 보입니다. 시뮬레이션 및 실제 환경에서의 평가 결과, WLA-0는 최첨단 수준의 다중 작업 및 장기 학습 능력을 달성하며, 예를 들어 RoboTwin2.0 Clean에서는 92.94%의 성공률, RMBench에서는 56.5%의 성공률을 보입니다. 또한, WLA-0는 액션 주석 없이 extit{다양한 로봇 비디오}로부터 새로운 작업을 직접 학습할 수 있는 잠재력을 가지고 있습니다.
We propose world-language-action (WLA) models as a new class of embodied foundation models. WLA takes textual instructions, images, and robot states as inputs to jointly predict textual subtasks, subgoal images, and robot actions, conjoining the \emph{world modeling interface} to learn from extensive egocentric videos as in the world-action model (WAM) and the \emph{language reasoning} capacities to solve complex long-horizon tasks as in vision-language-action (VLA) models. At the core of WLA lies an \emph{autoregressive (AR)} Transformer backbone, instead of a bidirectional diffusion Transformer as in WAMs, to predict the \emph{next state}, comprising the \emph{semantic-level} textual intention and complementary \emph{fine-grained} physical dynamics. The physical dynamics are supervised by the world modeling objective based on a dedicated World Expert, and are leveraged to ease the characterization of the state-action correlation for the Action Expert. WLA leverages meta-queries to make the world prediction \emph{implicitly} impact the action generation so that the former can be disabled during inference. The world prediction can also be activated to enable test-time scaling for improved robot control. Our WLA-0 prototype, with 2B active parameters, achieves 40 ms per inference on an NVIDIA RTX 5090. Evaluations across simulated and real-world environments demonstrate that WLA-0 achieves state-of-the-art multi-task and long-horizon learning abilities, e.g., 92.94\% success rate on RoboTwin2.0 Clean and 56.5\% success rate on RMBench. WLA-0 also holds the promise to learn novel tasks directly from \emph{cross-embodiment robot videos} without action annotations.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.