2606.05979v1 Jun 04, 2026 cs.RO

통합 세계 모델링, 언어 추론 및 행동 합성 위한 월드-언어-액션 모델

World-Language-Action Model for Unified World Modeling, Language Reasoning, and Action Synthesis

Pengfei Liu
Pengfei Liu
Citations: 528
h-index: 7
Siqi Kou
Siqi Kou
Citations: 293
h-index: 9
Zhijie Wei
Zhijie Wei
Citations: 0
h-index: 0
Xiaowu Xia
Xiaowu Xia
Citations: 1
h-index: 1
Zhijie Deng
Zhijie Deng
Citations: 88
h-index: 4
Yi Yang
Yi Yang
Citations: 721
h-index: 11
Zhihong Liu
Zhihong Liu
Citations: 70
h-index: 1
Yiyang Chen
Yiyang Chen
Citations: 8
h-index: 1
Yanzhe Hu
Yanzhe Hu
Citations: 8
h-index: 2
Jianbo Zhou
Jianbo Zhou
Citations: 2
h-index: 1
Bo Zhao
Bo Zhao
Citations: 1
h-index: 1
Xueqi Li
Xueqi Li
Citations: 15
h-index: 2

본 논문에서는 새로운 유형의 임베디드 기초 모델인 월드-언어-액션 (WLA) 모델을 제안합니다. WLA는 텍스트 지침, 이미지, 로봇 상태를 입력으로 받아 텍스트 하위 작업, 목표 이미지, 로봇 액션을 동시에 예측하며, 월드-액션 모델 (WAM)과 같이 광범위한 1인칭 비디오 데이터를 통해 학습하는 extit{세계 모델링 인터페이스}와 비전-언어-액션 (VLA) 모델처럼 복잡하고 장기적인 작업을 해결하기 위한 extit{언어 추론 능력}을 결합합니다. WLA의 핵심은 WAM에서 사용되는 양방향 디퓨전 트랜스포머 대신, extit{자기 회귀 (AR)} 트랜스포머 백본을 사용하여 extit{다음 상태}를 예측하는데, 이는 extit{의미 수준의} 텍스트 의도와 보완적인 extit{세밀한} 물리적 동역학으로 구성됩니다. 물리적 동역학은 전용 월드 전문가를 기반으로 하는 세계 모델링 목표에 의해 지도 학습되며, 액션 전문가가 상태-액션 상관 관계를 쉽게 파악할 수 있도록 활용됩니다. WLA는 메타 쿼리를 사용하여 세계 예측이 extit{암묵적으로} 행동 생성에 영향을 미치도록 하며, 이를 통해 추론 시 세계 예측 기능을 비활성화할 수 있습니다. 또한, 세계 예측을 활성화하여 테스트 시간에 확장성을 확보하고 로봇 제어를 향상시킬 수도 있습니다. 20억 개의 활성 파라미터를 가진 WLA-0 프로토타입은 NVIDIA RTX 5090에서 추론당 40ms의 성능을 보입니다. 시뮬레이션 및 실제 환경에서의 평가 결과, WLA-0는 최첨단 수준의 다중 작업 및 장기 학습 능력을 달성하며, 예를 들어 RoboTwin2.0 Clean에서는 92.94%의 성공률, RMBench에서는 56.5%의 성공률을 보입니다. 또한, WLA-0는 액션 주석 없이 extit{다양한 로봇 비디오}로부터 새로운 작업을 직접 학습할 수 있는 잠재력을 가지고 있습니다.

Original Abstract

We propose world-language-action (WLA) models as a new class of embodied foundation models. WLA takes textual instructions, images, and robot states as inputs to jointly predict textual subtasks, subgoal images, and robot actions, conjoining the \emph{world modeling interface} to learn from extensive egocentric videos as in the world-action model (WAM) and the \emph{language reasoning} capacities to solve complex long-horizon tasks as in vision-language-action (VLA) models. At the core of WLA lies an \emph{autoregressive (AR)} Transformer backbone, instead of a bidirectional diffusion Transformer as in WAMs, to predict the \emph{next state}, comprising the \emph{semantic-level} textual intention and complementary \emph{fine-grained} physical dynamics. The physical dynamics are supervised by the world modeling objective based on a dedicated World Expert, and are leveraged to ease the characterization of the state-action correlation for the Action Expert. WLA leverages meta-queries to make the world prediction \emph{implicitly} impact the action generation so that the former can be disabled during inference. The world prediction can also be activated to enable test-time scaling for improved robot control. Our WLA-0 prototype, with 2B active parameters, achieves 40 ms per inference on an NVIDIA RTX 5090. Evaluations across simulated and real-world environments demonstrate that WLA-0 achieves state-of-the-art multi-task and long-horizon learning abilities, e.g., 92.94\% success rate on RoboTwin2.0 Clean and 56.5\% success rate on RMBench. WLA-0 also holds the promise to learn novel tasks directly from \emph{cross-embodiment robot videos} without action annotations.

2 Citations
0 Influential
5.5 Altmetric
29.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!