PhoneWorld: 전화 사용 에이전트 환경의 확장
PhoneWorld: Scaling Phone-Use Agent Environments
전화 사용 에이전트에 있어 중요한 문제는 실제 모바일 사용 행태를 포괄하는 제어 가능하고 재현 가능한 환경을 대규모로 구축하기 어렵다는 것입니다. 기존의 모바일 에이전트 벤치마크는 평가 측면에서 상당한 발전을 이루었지만, 새로운 전화 사용 환경을 대량으로 구축할 수 있는 확장 가능한 방법을 직접적으로 제공하지 않습니다. 본 논문에서는 PhoneWorld를 소개합니다. PhoneWorld는 실제 GUI 트레이jectory와 스크린샷을 활용하여 제어 가능한 전화 사용 환경, 실행 가능한 작업, 자동 검증 시스템 및 학습 데이터를 생성하는 재사용 가능한 파이프라인입니다. PhoneWorld는 기존의 방식처럼 하나의 모바일 벤치마크를 일일이 구축하는 대신, 실제 트레이jectory를 분석하여 어떤 화면이 중요한지, 화면들이 어떻게 연결되는지, 어떤 상호작용이 환경 상태를 변화시키는지, 그리고 어떤 사용자 목표가 자동 검증을 허용하는지를 파악합니다. 이러한 정보를 바탕으로 읽기 전용 앱 콘텐츠와 변경 가능한 상태를 갖춘 실행 가능한 모의 Android 애플리케이션을 구축하고, 동일한 환경에서 실행 가능한 작업, 규칙 기반 검증 시스템 및 학습 데이터를 생성합니다. 현재 구현된 PhoneWorld는 16개 분야에 걸쳐 34개의 애플리케이션을 포괄하며, 검색, 브라우징, 쇼핑, 예약, 미디어, 소셜 인터랙션 등과 같은 일반적인 소비자 모바일 사용 행태를 포함합니다. 고정된 학습 예산을 기준으로, AndroidWorld 기반의 기본 모델에서 AndroidWorld 코퍼스에 있는 1만 단계를 PhoneWorld의 광범위한 지도 학습 데이터로 대체하면 네 가지 평가 지표 모두가 동시에 향상됩니다. 특히 HYMobileBench는 17.7점, AndroidControl은 6.0점, AndroidWorld는 14.7점, 그리고 PhoneWorld 자체는 52.5점이 상승합니다. 또한, PhoneWorld 지도 학습 데이터의 양을 늘리면 PhoneWorld 성능이 크게 향상되며, 고정된 PhoneWorld 예산 내에서 애플리케이션 적용 범위를 확장하면 더욱 큰 성과를 얻을 수 있다는 것을 확인했습니다. 종합적으로 볼 때, PhoneWorld는 개별 모바일 벤치마크 구축에 집중하는 방식에서 벗어나 전화 사용 환경 자체의 공급 규모를 확대하는 데 초점을 맞추고 있습니다.
A central bottleneck for phone-use agents is that controllable, reproducible environments covering real mobile behavior are hard to build at scale. Existing mobile-agent benchmarks have made important progress on evaluation, but they do not by themselves provide a scalable way to construct many new phone-use environments. We present PhoneWorld, a reusable pipeline that converts real GUI trajectories and screenshots into controllable phone-use environments, executable tasks, automatic verifiers, and training rollouts. Rather than hand-building one mobile benchmark at a time, PhoneWorld uses real trajectories to recover which screens matter, how screens connect, which interactions must change environment state, and which user goals admit automatic verification. From these signals, it builds runnable mock Android apps backed by read-only app content and mutable state, then derives executable tasks, rule-based verifiers, and training rollouts from the same environments. In its current instantiation, PhoneWorld covers 34 apps across 16 domains, spanning common consumer mobile behaviors such as search, browsing, shopping, booking, media, and social interaction. Under a fixed training budget, replacing 10K steps from an auxiliary AndroidWorld corpus in an AndroidWorld-based baseline with broad PhoneWorld supervision improves all four evaluation benchmarks at once, raising HYMobileBench by 17.7 points, AndroidControl by 6.0 points, AndroidWorld by 14.7 points, and PhoneWorld by 52.5 points. We then study two additional scaling questions: increasing the amount of PhoneWorld supervision strongly improves PhoneWorld performance, and under a fixed PhoneWorld budget, expanding app coverage yields even larger gains. Overall, PhoneWorld shifts the focus from building one mobile benchmark at a time to scaling the supply of phone-use environments themselves.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.