Act2Intention: GUI 액션을 통해 사용자 의도를 추론하여 능동적인 모바일 에이전트를 개발하기 위한 벤치마크
Act2Intention: A Benchmark For Developing Active Mobile Agents Through Inferring User Intention from GUI Actions
다중 모달 대규모 언어 모델(MLLM)을 기반으로 하는 모바일 GUI 에이전트는 인간-컴퓨터 지능 분야에서 잠재력을 보여줍니다. 그러나 현재 연구는 주로 반응적인 작업 실행에 집중하고 있으며, 능동적인 에이전트의 핵심 요구 사항인 사용자 의도에 대한 종합적인 이해-예측-실행 과정을 부족하게 다루고 있습니다. 본 논문에서는 이해, 사용자 의도 예측 및 의사 결정 실행을 통합하여 능동적인 모바일 에이전트를 구축하는 Act2Intention 프레임워크를 제안합니다. 먼저 데이터 수집 및 검증된 생성 방법을 통해 72,511개의 의도와 52개의 앱에 걸쳐 70만 건 이상의 액션으로 구성된 Act2Intention 벤치를 구축하여 연속적인 의도-액션 경로를 통한 능동적인 에이전트 평가를 위한 최초의 벤치마크를 제시합니다. 또한, Act2Intention Agent를 개발하여 의도 이해, 개인화된 의도 예측 및 경험 기반의 의사 결정 실행을 통해 능동적인 서비스를 제공합니다. 실험 결과는 Act2Intention 벤치를 사용한 지도 학습 미세 조정이 동일한 에이전트 프레임워크 내에서 의도 이해, 예측 및 실행에 대해 각각 +32.0 Acc-S, +10.25 Acc-S, +6.9 SSR의 절대적인 성능 향상을 가져왔음을 보여줍니다. 이러한 성공은 Act2Intention 벤치의 필요성과 가치를 강조하며, 능동적인 에이전트 개발 및 평가를 위한 표준화된 플랫폼을 구축하고 결과적으로 의도 기반 인간-컴퓨터 상호 작용 연구의 길을 열어줍니다.
Mobile GUI Agents powered by multimodal large language models (MLLMs) show promise in human-computer intelligence. However, current research primarily focuses on reactive task execution while lacking a comprehensive understanding-prediction-execution process for user intentions, which are the core requirements of active agents. In this paper, we propose the Act2Intention framework that builds an active mobile agent by integrating understanding, predicting user intentions, and executing decisions. First, we construct the Act2Intention Bench through data collection and validated generation, comprising 72,511 intentions and over 700,000 actions across 52 apps, thereby establishing the first benchmark for evaluating proactive agents via continuous intention-action trajectories. We further develop the Act2Intention Agent, achieving proactive services through Proactive-oriented Intention Understanding, Personalized Proactive Intention Prediction, and Experience-guided Intention Execution. Experimental results show that supervised fine-tuning on Act2Intention Bench yields absolute improvements of +32.0 Acc-S, +10.25 Acc-S, and +6.9 SSR points over non-fine-tuned counterparts under the same agent framework for intention understanding, prediction, and execution, respectively. This success underscores the necessity and value of the Act2Intention Bench, which establishes a standardized platform for developing and evaluating proactive agents and consequently paves the way for research on intention-driven human-computer interaction.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.