2608.13489v1 Aug 13, 2026 cs.CV

DreamX-Phi 1.0: 로봇 조작을 위한 액션 기반 비디오 월드 모델

DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Pengfei Zhang
Pengfei Zhang
Citations: 64
h-index: 3
Xiangxiang Chu
Xiangxiang Chu
Citations: 292
h-index: 10
Jing Tang
Jing Tang
Citations: 181
h-index: 5
Dream Team
Dream Team
Citations: 0
h-index: 0
Rui Chen
Rui Chen
AMAP, Alibaba Group
Citations: 241
h-index: 6
Ge Li
Ge Li
Citations: 46
h-index: 1
Qingfeng Shi
Qingfeng Shi
Citations: 3
h-index: 1
Datao Tang
Datao Tang
Citations: 181
h-index: 4

본 논문에서는 로봇 조작을 위한 액션 기반 비디오 월드 모델인 **DreamX-Phi 1.0**을 제시합니다. 이 모델은 관찰된 프레임, 언어 지시 및 엔드 이펙터 자세와 그리퍼 상태로 구성된 미리 정의된 행동 시퀀스를 입력으로 받아 결과적으로 나타날 미래의 관찰 결과를 예측합니다. 그러나 현실성만으로는 정확성을 보장할 수 없습니다. 설득력 있는 예측이라도 잘못된 팔을 움직이거나 조작 대상 물체를 잃을 수 있습니다. 각 팔이 지정된 경로를 준수하도록 하기 위해, **PRoPE 방식의 기하학적 인코딩**을 통해 어텐션에 각 팔의 $ ext{SE}(3)$ 변환을 주입하여 팔의 동일성을 유지하고 강체 운동 구조를 보존합니다. 액션 제어만으로는 장면의 기하학적 구조나 작은 조작 대상 물체의 변화를 완전히 제한할 수 없으므로, 장면 수준의 기하학적 정보를 처리하기 위한 경량화된 **깊이 분기(depth branch)**를 추가하고, 고정된 **V-JEPA 모델**을 사용한 **SAM3 마스크**를 통해 조작 과정 전반에 걸쳐 물체의 일관성을 유지합니다. 또한, 다단계 생성기를 효율적인 배포를 위해 몇 단계의 학생 모델로 증류(distillation)했습니다. 논문 작성 시점 기준으로, 저희 모델은 WorldArena~2.0 Challenge의 Track~1에서 1위를 차지하고 Track~2에서 2위를 차지했습니다. 저희 모델과 코드는 공개적으로 제공될 예정입니다.

Original Abstract

We present \textbf{DreamX-Phi 1.0}, an action-conditioned video world model for robotic manipulation that, given an observed frame, a language instruction, and a prescribed action sequence comprising end-effector poses and gripper states, predicts the resulting future observations. Yet realism alone does not guarantee faithfulness: a convincing rollout can still move the wrong arm or lose the manipulated object. To ensure the prediction respects each arm's commanded path, we inject per-arm $\mathrm{SE}(3)$ transformations into attention via \textbf{PRoPE-style geometric encoding}, preserving arm identity and rigid-motion structure. Action control alone does not fully constrain scene geometry or the evolution of small manipulated objects. We therefore add a lightweight \textbf{depth branch} for scene-level geometry and use \textbf{SAM3 masks} with a frozen \textbf{V-JEPA teacher} to maintain object consistency throughout grasping. We further distill the multi-step generator into a few-step student via distribution-matching distillation for efficient deployment. At the time of writing, \model{} achieves first place on Track~1 and second place on Track~2 of the WorldArena~2.0 Challenge. Our model and code will be publicly available.

0 Citations
0 Influential
5 Altmetric
25.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!