2606.29705v1 Jun 29, 2026 cs.AI

GUICrafter: 대규모 비주석 스크린샷을 활용한 약하게 지도된 GUI 에이전트

GUICrafter: Weakly-Supervised GUI Agent Leveraging Massive Unannotated Screenshots

Runqiu Yin
Runqiu Yin
Citations: 536
h-index: 1
Meng-Hao Guo
Meng-Hao Guo
Citations: 166
h-index: 6
Shi-Min Hu
Shi-Min Hu
Citations: 204
h-index: 7
Sunqi Fan
Sunqi Fan
Citations: 13
h-index: 2
Yongming Rao
Yongming Rao
Citations: 53
h-index: 3
Lingshan Chen
Lingshan Chen
Citations: 0
h-index: 0
Qingle Liu
Qingle Liu
Citations: 0
h-index: 0

데이터는 현대 인공지능의 근간으로서, 현재 기초 모델 개발에 큰 영향을 미쳐왔습니다. 자연스럽게 연구자들은 이 패러다임을 GUI 에이전트 분야로 확장하여 강력한 GUI 에이전트를 구축하고자 합니다. 그러나 GUI 에이전트 데이터는 인터넷에서 직접 수집하기 어려워, 대규모로 수집하는 데 비용과 노력이 많이 듭니다. 그 결과, 현재의 GUI 에이전트는 장치 간 일반화 성능이 낮고, 세밀한 GUI 요소에 대한 시각적 이해 능력이 제한적인 문제가 있습니다. 이러한 데이터 문제를 해결하고자, 본 연구에서는 대규모 비주석 스크린샷을 활용하여 인간 주석 의존도를 크게 줄이는 약하게 지도된 GUI 에이전트인 GUICrafter를 제안합니다. GUICrafter는 두 단계로 진행되는 커리큘럼 학습 프레임워크를 통해 GUI 에이전트를 훈련합니다. 먼저, 모델은 대규모 비주석 스크린샷 및 웹페이지로부터 시각적 정보를 학습하며, 인간 주석 없이 GUI 상호작용에 내재된 풍부한 맥락 신호를 활용합니다. 두 번째 단계에서는 소량의 고품질 데이터를 사용하여 강화 학습을 통해 모델을 보정합니다. 실험 결과, GUICrafter는 UI-TARS와 같은 고급 시스템과 경쟁력 있는 성능, 또는 더 나은 성능을 달성하며, 데이터 사용량은 0.1%에 불과했습니다. 또한 동일한 양의 주석 데이터를 사용할 때, GUI-R1을 포함한 이전 방법보다 우수한 성능을 보였습니다. 코드, 데이터 및 모델은 https://github.com/fansunqi/GUICrafter 에서 확인할 수 있습니다.

Original Abstract

Data, as the fundamental substrate of modern intelligence, has greatly driven the development of current foundation models. Naturally, researchers aim to extend this paradigm to the domain of GUI agents, hoping to build strong GUI agents through a similar paradigm. However, GUI agent data cannot be directly harvested from the internet, making it costly and difficult to collect at scale. As a result, current GUI agents suffer from poor cross-device generalization and limited visual grounding ability for fine-grained GUI elements. As an attempt to address data challenge in GUI agents, we propose GUICrafter, a weakly-supervised GUI agent leveraging massive unannotated screenshots to substantially reduce the reliance on expensive human annotations. GUICrafter explores a curriculum learning framework for training GUI agents through two progressive stages. First, the model learns visual grounding from large-scale unannotated screenshots and webpages, leveraging the rich contextual signals inherent in GUI interactions without human annotations. Then, in Stage 2, we leverage a small amount of high-quality data to calibrate the model via reinforcement learning. Experiments show that GUICrafter achieves competitive, or even superior, performance to advanced systems like UI-TARS while using only 0.1% of its data. Furthermore, under the same amount of annotated data, GUICrafter surpasses all previous methods such as GUI-R1. Code, data, and models are available at https://github.com/fansunqi/GUICrafter.

0 Citations
0 Influential
28.993061443341 Altmetric
0.0 Score
Original PDF
2

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!