모바일 GUI 탐색을 위한 시각-언어 에이전트의 확장, 벤치마킹 및 추론 연구
Scaling, Benchmarking, and Reasoning of Vision-Language Agents for Mobile GUI Navigation
시각-언어 모델(VLM)은 모바일 GUI 탐색 분야에서 빠른 발전을 보여주고 있습니다. 본 논문에서는 이 분야의 VLM 기반 에이전트에 대한 데이터 규모 확장, 벤치마킹 및 추론에 대한 체계적인 연구를 제시합니다. 엄밀한 평가를 위해, 실제 650개 이상의 중국 모바일 애플리케이션에서 추출된 16,000개 이상의 실세계 작업을 포함하는 대규모 데이터셋인 HyperTrack과 VLM의 오프라인 GUI 탐색 작업에 대한 통합 벤치마킹 도구인 GUIEvalKit을 소개합니다. HyperTrack을 사용하여 지도 학습 및 강화 학습 기반 미세 조정을 모두 사용했을 때, 학습 데이터 규모가 성능에 미치는 영향을 분석했습니다. 결과는 강화 학습 기반 미세 조정이 지도 학습 기반 미세 조정보다 일관되게 우수한 성능을 보이며, 특히 일반화 성능이 뛰어나다는 것을 보여주며, 이는 데이터 규모 확장과 강화 학습 간의 시너지 효과를 강조합니다. 또한 GUIEvalKit을 활용하여 최첨단(SOTA) VLM을 벤치마킹하고, 상호 작용 기록 및 추론 능력이 작업 완료에 미치는 영향을 분석했습니다. HyperTrack과 GUIEvalKit은 모바일 GUI 탐색 작업에서 VLM 에이전트를 개발하고 평가하기 위한 포괄적인 플랫폼을 제공합니다.
Vision-Language Models (VLMs) have shown rapid progress in mobile GUI navigation. This paper presents a systematic study of data scaling, benchmarking, and reasoning for VLM-based agents in this domain. To facilitate rigorous evaluation, we introduce HyperTrack, a large-scale dataset with over 16000 real-world tasks across more than 650 Chinese mobile applications, along with GUIEvalKit, an open-source toolkit for unified benchmarking of VLMs on offline GUI navigation tasks. Using HyperTrack, we analyze the effects of training data scale on both supervised and reinforcement-based finetuning. Our results show that reinforcement-based finetuning consistently outperforms supervised finetuning, particularly in out-of-domain settings, highlighting the synergy between data scaling and reinforcement learning. Leveraging GUIEvalKit, we further benchmark state-of-the-art (SOTA) VLMs and analyze how interaction history and reasoning capabilities influence task completion. Together, HyperTrack and GUIEvalKit provide a comprehensive platform for developing and evaluating VLM agents in mobile GUI navigation tasks.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.