DreamX-Phi 1.0: 로봇 조작을 위한 액션 기반 비디오 월드 모델
DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation
본 논문에서는 로봇 조작을 위한 액션 기반 비디오 월드 모델인 **DreamX-Phi 1.0**을 제시합니다. 이 모델은 관찰된 프레임, 언어 지시 및 엔드 이펙터 자세와 그리퍼 상태로 구성된 미리 정의된 행동 시퀀스를 입력으로 받아 결과적으로 나타날 미래의 관찰 결과를 예측합니다. 그러나 현실성만으로는 정확성을 보장할 수 없습니다. 설득력 있는 예측이라도 잘못된 팔을 움직이거나 조작 대상 물체를 잃을 수 있습니다. 각 팔이 지정된 경로를 준수하도록 하기 위해, **PRoPE 방식의 기하학적 인코딩**을 통해 어텐션에 각 팔의 $ ext{SE}(3)$ 변환을 주입하여 팔의 동일성을 유지하고 강체 운동 구조를 보존합니다. 액션 제어만으로는 장면의 기하학적 구조나 작은 조작 대상 물체의 변화를 완전히 제한할 수 없으므로, 장면 수준의 기하학적 정보를 처리하기 위한 경량화된 **깊이 분기(depth branch)**를 추가하고, 고정된 **V-JEPA 모델**을 사용한 **SAM3 마스크**를 통해 조작 과정 전반에 걸쳐 물체의 일관성을 유지합니다. 또한, 다단계 생성기를 효율적인 배포를 위해 몇 단계의 학생 모델로 증류(distillation)했습니다. 논문 작성 시점 기준으로, 저희 모델은 WorldArena~2.0 Challenge의 Track~1에서 1위를 차지하고 Track~2에서 2위를 차지했습니다. 저희 모델과 코드는 공개적으로 제공될 예정입니다.
We present \textbf{DreamX-Phi 1.0}, an action-conditioned video world model for robotic manipulation that, given an observed frame, a language instruction, and a prescribed action sequence comprising end-effector poses and gripper states, predicts the resulting future observations. Yet realism alone does not guarantee faithfulness: a convincing rollout can still move the wrong arm or lose the manipulated object. To ensure the prediction respects each arm's commanded path, we inject per-arm $\mathrm{SE}(3)$ transformations into attention via \textbf{PRoPE-style geometric encoding}, preserving arm identity and rigid-motion structure. Action control alone does not fully constrain scene geometry or the evolution of small manipulated objects. We therefore add a lightweight \textbf{depth branch} for scene-level geometry and use \textbf{SAM3 masks} with a frozen \textbf{V-JEPA teacher} to maintain object consistency throughout grasping. We further distill the multi-step generator into a few-step student via distribution-matching distillation for efficient deployment. At the time of writing, \model{} achieves first place on Track~1 and second place on Track~2 of the WorldArena~2.0 Challenge. Our model and code will be publicly available.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.