2606.31410v1 Jun 30, 2026 cs.AI

Xiaomi-GUI-0 기술 보고서

Xiaomi-GUI-0 Technical Report

Pei Fu
Pei Fu
Citations: 28
h-index: 4
Ruoceng Zhang
Ruoceng Zhang
Citations: 9
h-index: 2
Shaojie Zhang
Shaojie Zhang
Citations: 9
h-index: 2
Jiahui Yang
Jiahui Yang
Citations: 126
h-index: 3
Zhenbo Luo
Zhenbo Luo
Citations: 213
h-index: 7
Jian Luan
Jian Luan
Citations: 947
h-index: 13
Niu Lian
Niu Lian
Citations: 42
h-index: 4
Hengxu Qu
Hengxu Qu
Citations: 303
h-index: 5
Qinzhuo Wu
Qinzhuo Wu
Citations: 118
h-index: 4
Pengzhi Gao
Pengzhi Gao
Citations: 72
h-index: 5
Wei Liu
Wei Liu
Citations: 495
h-index: 12
Tong-I Chen
Tong-I Chen
Citations: 8
h-index: 2
Anan Du
Anan Du
Citations: 20
h-index: 3
Yike Liu
Yike Liu
Citations: 19
h-index: 2
Wanxia Cao
Wanxia Cao
Citations: 10
h-index: 1
Chengzhen Duan
Chengzhen Duan
Citations: 0
h-index: 0
Zhehao Yu
Zhehao Yu
Citations: 0
h-index: 0
Shiqi Cui
Shiqi Cui
Citations: 5
h-index: 1
Shukai Jia
Shukai Jia
Citations: 0
h-index: 0
Wen‐Tza Lu
Wen‐Tza Lu
Citations: 1
h-index: 1
Jiatong Sun
Jiatong Sun
Citations: 0
h-index: 0
Chengkai Tan
Chengkai Tan
Citations: 0
h-index: 0
Tao Xiong
Tao Xiong
Citations: 401
h-index: 8
Jian-Hao Zhu
Jian-Hao Zhu
Citations: 0
h-index: 0
Congyu Zou
Congyu Zou
Citations: 0
h-index: 0
Fazhan Liu
Fazhan Liu
Citations: 0
h-index: 0
Hui Liu
Hui Liu
Citations: 227
h-index: 4
Yuanfa Li
Yuanfa Li
Citations: 0
h-index: 0
Haoyuan Sun
Haoyuan Sun
Citations: 108
h-index: 7
Yajie Wang
Yajie Wang
Citations: 0
h-index: 0
Changqiao Wu
Changqiao Wu
Citations: 338
h-index: 4
Yu Yuan
Yu Yuan
Citations: 0
h-index: 0

그래픽 사용자 인터페이스(GUI) 에이전트는 시각-언어 모델을 기반으로 하며, 실제 애플리케이션에서 터치, 스와이프, 텍스트 입력 및 탐색과 같은 인터페이스 액션을 통해 사용자의 작업을 처음부터 끝까지 수행합니다. 그러나 기존 GUI 에이전트는 주로 오프라인 트래젝토리, 시뮬레이션 환경 및 표준화된 벤치마크를 사용하여 학습되고 평가됩니다. 이러한 요소들은 인터페이스 레이아웃, 상호 작용 로직 및 비정상 상태 분포 측면에서 실제 애플리케이션과 크게 다르며, 계정 상태, 권한 대화, 결제 인증 및 위험 관리와 같이 지속적으로 상태 분포를 변화시키는 실제 환경에서의 실행 안정성을 정확하게 반영하지 못합니다. 이러한 격차를 줄이기 위해, 우리는 실제 모바일 환경을 위한 네이티브 멀티모달 GUI 에이전트인 Xiaomi-GUI-0을 제안하며, 이는 실제 장치 폐쇄 루프 내에서 학습되고 평가됩니다. 핵심은 물리적 장치가 주요 실행 환경이고 샌드박스가 보조 지원 역할을 제공하는 실제 장치 중심의 하이브리드 인프라입니다. 이를 통해 데이터 수집, 학습, 배포 및 평가는 실제 배포와 유사한 실행 분포를 공유합니다. 우리는 고빈도 주요 작업, 일반화 능력을 향상시키는 희소 데이터 및 반성 및 기억 능력 향상을 위한 데이터를 포함하는 다중 소스 훈련 데이터를 구축하고, 실패 트래젝토리를 수정된 액션, 반성적인 설명 및 복구 데모로 변환하는 오류 기반 데이터 순환 시스템을 도입했습니다. 모델은 지도 학습 파인튜닝, 단계별 강화 학습 및 에이전트 중심 강화 학습이라는 점진적인 세 단계 파이프라인을 통해 훈련됩니다. 공개 벤치마크와 자체 개발한 RealMobile에서 평가된 결과, Xiaomi-GUI-0은 RealMobile에서 72.0%의 성공률과 AndroidWorld에서 78.9%의 성공률을 달성했으며, 실제 작업에서의 실행 안정성과 비정상 상태 인식 능력을 크게 향상시켰습니다.

Original Abstract

Graphical user interface (GUI) agents build on vision-language models to complete user tasks end-to-end in real applications through interface actions such as tapping, swiping, text entry, and navigation. However, existing GUI agents are trained and evaluated largely on offline trajectories, simulated environments, and standardized benchmarks. These differ substantially from real applications in interface layout, interaction logic, and abnormal-state distribution, and cannot faithfully characterize execution stability in real-world use, where account states, permission dialogs, payment authentication, and risk control continually reshape the state distribution and open a persistent gap between benchmark scores and real usability. To close this gap, we propose Xiaomi-GUI-0, a native multimodal GUI agent for real mobile environments, trained and evaluated within a real-device closed loop. At its core is a real-device-dominant hybrid infrastructure, where physical devices are the primary execution environment and sandboxes provide auxiliary support, so that data collection, training, rollout, and evaluation share an execution distribution close to real deployment. We construct multi-source training data spanning high-frequency head tasks, high-generalization data for long-tail intents, and capability-enhancement data for reflection and memory, and introduce an error-driven data flywheel that turns failure trajectories into corrected actions, reflective explanations, and recovery demonstrations. The model is trained through a progressive three-stage pipeline of supervised fine-tuning, step-level reinforcement learning, and agentic reinforcement learning. Evaluated on public benchmarks and our in-house RealMobile, Xiaomi-GUI-0 achieves 72.0% success on RealMobile and 78.9% on AndroidWorld, while substantially improving execution stability and abnormal-state recognition in real-world tasks.

1 Citations
0 Influential
6.5 Altmetric
33.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!