Qwen-UI-Agent 기술 보고서: 차세대 실세계 중심 기반 그래픽 사용자 인터페이스 에이전트 개발을 향하여
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents
그래픽 사용자 인터페이스(GUI) 에이전트는 기존 디지털 장치에 대한 범용 실행기로 발전할 잠재력을 가지고 있습니다. 이러한 에이전트를 실제 사용 환경으로 확장하기 위해, 우리는 실제 장치에서 안정적으로 작동하고, 다양한 플랫폼에서 워크플로우를 실행하며, GUI 상호 작용과 CLI(명령줄 인터페이스) 실행을 결합하고, 장기적인 작업을 수행하며, 유용한 서비스를 능동적으로 제공하고, 최소한의 인간 개입으로 자체 기능을 지속적으로 개선하는 에이전트를 구상합니다. 이러한 비전을 바탕으로, 우리는 모바일, 컴퓨터 사용, 웹 및 DeepSearch 환경을 포괄하는 실세계 중심 기반 GUI 에이전트인 Qwen-UI-Agent를 소개합니다. Qwen-UI-Agent는 다양한 샌드박스 환경과 대규모 실제 장치 모바일 런타임을 결합하여, 통일된 동작 공간을 통해 GUI 작업과 CLI 실행을 통합하고, 단일 모델 턴으로 일괄 작업을 생성합니다. AutoResearch 스타일의 데이터 플라이휠은 에이전트를 사용하여 작업 및 환경을 구축하고, 오류를 진단하며, 후속 반복을 계획합니다. 온라인 강화 학습(RL)을 통해 100회 이상의 트랙토리에 대한 학습을 지원하며, 1만 개 이상의 동시 환경을 활용하여 빠른 배포를 가능하게 합니다. 경량화된 하니스 레이어는 모바일 및 컴퓨터 환경에서 능동적인 서비스 시작과 상태 기반 워크플로우를 지원합니다. 다양한 평가 결과, Qwen-UI-Agent는 모바일 사용 벤치마크에서 최고 수준의 성능을 달성했으며, Opus 4.8, Gemini 3.1 Pro 및 GPT-5.6 Sol과 같은 최첨단 모델과의 경쟁에서도 우수한 성능을 보였습니다. 모바일 환경에서는 MobileWorld에서 82.1%, MobileWorld-Real에서 92.2%, AndroidDaily에서 97.5%의 정확도를 달성했습니다. 컴퓨터 사용 환경에서는 OSWorld-Verified에서 79.5%의 정확도를, OSWorld-v2에서는 40.0%의 부분적인 진행률을 보였습니다. 웹 사용 및 GUI 이해 측면에서는 WebArena에서 73.6%, ScreenSpot-Pro에서 81.5%의 정확도를 달성했습니다.
GUI agents have the potential to become a general purpose executor over existing digital devices. To advance them toward real-world use, we envision agents that operate reliably on real devices, execute workflows across platforms, combine GUI interaction with CLI execution, complete long-horizon tasks, proactively initiate useful services, and autonomously improve their capabilities with minimal human effort. Guided by this vision, we present Qwen-UI-Agent, a real-world centric foundation GUI agent spanning mobile, computer-use, web, and DeepSearch environments. Qwen-UI-Agent combines diverse sandbox environments with a large-scale real-device mobile runtime. Its unified action space interleaves GUI operations with CLI execution and generates batched actions in a single model turn. An AutoResearch-style data flywheel uses agents to construct tasks and environments, diagnose failures, and plan subsequent iterations. Online RL supports training on trajectories exceeding 100 turns, with over 10,000 concurrent environments accelerating rollout. A lightweight harness layer supports proactive service initiation and stateful workflows across mobile and computer. Across a broad suite of evaluations, Qwen-UI-Agent sets state-of-the-art performance on mobile-use benchmarks while delivering competitive performance on computer- and browser-use tasks against frontier models, including Opus 4.8, Gemini 3.1 Pro, and GPT-5.6 Sol. On mobile use, it achieves 82.1% on MobileWorld, 92.2% on MobileWorld-Real, and 97.5% on AndroidDaily. On computer use, it achieves 79.5% on OSWorld-Verified and a 40.0% partial-progress score on OSWorld-v2. On browser use and GUI grounding, it achieves 73.6% on WebArena and 81.5% on ScreenSpot-Pro, respectively.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.