2607.11643v1 Jul 13, 2026 cs.RO

Xiaomi-Robotics-U0: 월드 기반 모델을 활용한 통합된 로봇 에이전트 생성

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model

Han Zhao
Han Zhao
Citations: 957
h-index: 16
Jiahang Cao
Jiahang Cao
Citations: 40
h-index: 3
Hangjun Ye
Hangjun Ye
Citations: 282
h-index: 8
Hengxu Qu
Hengxu Qu
Citations: 303
h-index: 5
Hongyu Yan
Hongyu Yan
Citations: 6
h-index: 2
Qiwei Li
Qiwei Li
Citations: 126
h-index: 6
Jun Guo
Jun Guo
Citations: 191
h-index: 6
Nan Sun
Nan Sun
Citations: 18
h-index: 2
Xinghang Li
Xinghang Li
Citations: 1,111
h-index: 8
Huaping Liu
Huaping Liu
Citations: 163
h-index: 4
Jason Li
Jason Li
Citations: 18
h-index: 2
Yunhong Wang
Yunhong Wang
Citations: 38
h-index: 3
Long Qian
Long Qian
Citations: 471
h-index: 6
Hang Lai
Hang Lai
Citations: 7
h-index: 1
Yueze Wang
Yueze Wang
Citations: 2,740
h-index: 16
Xi Chen
Xi Chen
Citations: 0
h-index: 0
Jingen Qu
Jingen Qu
Citations: 14
h-index: 2
Jia Song
Jia Song
Citations: 0
h-index: 0
Futeng Liu
Futeng Liu
Citations: 9
h-index: 1
Wanli Peng
Wanli Peng
Citations: 0
h-index: 0
Heyun Wang
Heyun Wang
Citations: 0
h-index: 0
Caoyu Xia
Caoyu Xia
Citations: 0
h-index: 0
Jack Zhao
Jack Zhao
Citations: 0
h-index: 0
Diyun Xiang
Diyun Xiang
Citations: 57
h-index: 4

최근의 기초 이미지 및 비디오 생성 모델은 강력한 일반화 능력과 제어 기능을 제공하지만, 다중 시점 일관성, 기하학적 응집성, 그리고 로봇 특성을 고려해야 하는 실제 환경에서의 활용에는 제한이 있습니다. 기존 방법들은 주로 제한된 로봇 데이터를 사용하여 기초 모델을 조정하며, 이 과정에서 대규모 사전 훈련을 통해 얻은 시각적 지식이 손실되는 경우가 많습니다. 본 논문에서는 통합된 로봇 에이전트 생성을 위한 380억 개의 파라미터를 가진 다중 모드 오토 회귀 모델인 Xiaomi-Robotics-U0을 제안합니다. 이는 로봇 기반의 생성 과정을 기초 이미지 및 비디오 생성의 확장으로 보고, 텍스트-이미지 생성, 이미지 편집, 로봇 환경 시나리오 생성, 에이전트 전송(transfer), 그리고 로봇 비디오 생성을 공동 최적화합니다. 이러한 통합 프레임워크는 사전 훈련된 월드 기반 모델의 일반화 능력을 유지하면서 로봇 환경에 맞게 적용할 수 있도록 합니다. Xiaomi-Robotics-U0은 여러 로봇 환경에서 고품질의 다중 시점 장면 생성을 지원하는 최초의 모델이며, 다중 시점 일관성과 상호 작용 역학을 유지하면서 세밀한 편집 기능을 위한 구조화되고 제어 가능한 에이전트 전송 방식을 도입합니다. 이 모델은 단일 단계 및 순차적 생성 작업에서 최첨단 성능을 달성했으며, 로봇 환경 시나리오 생성 및 전송에 대한 인간 평가에서 GPT-Image-2.0보다 뛰어난 성능을 보였습니다. 또한 World Arena에서 로봇 비디오 생성을 위한 1위를 차지했으며, pi_0.5의 일반화되지 않은 데이터(out-of-distribution) 작업 성공률을 36.9%에서 63.2%로 향상시켰습니다. 이러한 결과는 기초 월드 모델이 로봇 환경 모델이자 로봇 지능 개발을 위한 확장 가능한 데이터 엔진으로 활용될 수 있음을 보여줍니다. 코드 및 체크포인트는 https://robotics.xiaomi.com/xiaomi-robotics-u0.html 에서 확인할 수 있습니다.

Original Abstract

Recent foundation image and video generation models offer strong generalization and controllability, but their direct application to embodied scenarios is limited by requirements for multi-view consistency, geometric coherence, and robot embodiment constraints. Existing methods typically adapt foundation models with limited robot data, often sacrificing visual knowledge acquired during large-scale pre-training. We present Xiaomi-Robotics-U0, a 38-billion-parameter multimodal autoregressive model for unified embodied synthesis. It treats embodied generation as an extension of foundation image and video generation and jointly optimizes text-to-image generation, image editing, embodied scene generation, embodied transfer, and embodied video generation. This unified framework preserves the generalization of the pre-trained world foundation model while adapting it to embodied settings. Xiaomi-Robotics-U0 is the first model to support high-quality multi-view scene generation across multiple robot embodiments and to introduce structured, controllable embodied transfer for fine-grained editing while preserving multi-view consistency and interaction dynamics. It achieves state-of-the-art results on single-step and sequential generation tasks, outperforming GPT-Image-2.0 in human evaluations of embodied scene generation and transfer, ranking first on World Arena for embodied video generation, and improving the out-of-distribution success rate of pi_0.5 from 36.9% to 63.2% on challenging real-world manipulation tasks. These results show that foundation world models can serve both as embodied world models and scalable data engines for embodied intelligence. Code and checkpoints are available at https://robotics.xiaomi.com/xiaomi-robotics-u0.html.

1 Citations
0 Influential
8 Altmetric
41.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!