PSPA-Bench: 스마트폰 GUI 에이전트를 위한 개인 맞춤형 벤치마크
PSPA-Bench: A Personalized Benchmark for Smartphone GUI Agent
스마트폰 GUI 에이전트는 앱 인터페이스를 직접 조작하여 작업을 수행하며, 깊은 시스템 통합 없이도 다양한 기능을 제공할 수 있습니다. 그러나 실제 스마트폰 사용은 매우 개인화되어 있으며, 사용자는 다양한 워크플로우와 선호도를 가지므로, 에이전트는 일반적인 솔루션이 아닌 맞춤형 지원을 제공해야 합니다. 기존 GUI 에이전트 벤치마크는 사용자별 데이터 부족과 세분화된 평가 지표의 부재로 인해 이러한 개인화 측면을 충분히 반영하지 못합니다. 이러한 격차를 해소하기 위해, 스마트폰 GUI 에이전트의 개인화 기능을 평가하는 데 특화된 벤치마크인 PSPA-Bench를 제안합니다. PSPA-Bench는 10가지 대표적인 일상 사용 시나리오와 22개의 모바일 앱에 걸쳐 수집된 12,855개 이상의 개인화된 명령어를 포함하며, 구조를 고려한 프로세스 평가 방법을 통해 에이전트의 개인화된 기능을 세분화된 수준에서 측정합니다. PSPA-Bench를 사용하여 11개의 최첨단 GUI 에이전트를 벤치마킹했습니다. 결과는 현재 방법이 개인화된 환경에서 성능이 저조하며, 가장 강력한 에이전트조차 제한적인 성공만을 거둔다는 것을 보여줍니다. 분석 결과, 개인화된 GUI 에이전트 발전을 위한 세 가지 방향을 제시합니다. (1) 추론 기반 모델이 일반적인 LLM보다 우수한 성능을 보이며, (2) 인지 능력은 단순하지만 중요한 기능이며, (3) 반사 및 장기 기억 메커니즘은 적응 능력을 향상시키는 데 핵심적인 역할을 합니다. 이러한 결과들을 종합적으로 고려할 때, PSPA-Bench는 개인화된 GUI 에이전트에 대한 체계적인 연구와 미래 발전을 위한 기반을 제공합니다.
Smartphone GUI agents execute tasks by operating directly on app interfaces, offering a path to broad capability without deep system integration. However, real-world smartphone use is highly personalized: users adopt diverse workflows and preferences, challenging agents to deliver customized assistance rather than generic solutions. Existing GUI agent benchmarks cannot adequately capture this personalization dimension due to sparse user-specific data and the lack of fine-grained evaluation metrics. To address this gap, we present PSPA-Bench, the benchmark dedicated to evaluating personalization in smartphone GUI agents. PSPA-Bench comprises over 12,855 personalized instructions aligned with real-world user behaviors across 10 representative daily-use scenarios and 22 mobile apps, and introduces a structure-aware process evaluation method that measures agents' personalized capabilities at a fine-grained level. Through PSPA-Bench, we benchmark 11 state-of-the-art GUI agents. Results reveal that current methods perform poorly under personalized settings, with even the strongest agent achieving limited success. Our analysis further highlights three directions for advancing personalized GUI agents: (1) reasoning-oriented models consistently outperform general LLMs, (2) perception remains a simple yet critical capability, and (3) reflection and long-term memory mechanisms are key to improving adaptation. Together, these findings establish PSPA-Bench as a foundation for systematic study and future progress in personalized GUI agents.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.