작업 계획을 위한 자율적인 경험 탐색 및 후향적 경험 활용을 통한 GUI 에이전트의 역량 강화
Empowering GUI Agents via Autonomous Experience Exploration and Hindsight Experience Utilization for Task Planning
멀티모달 웹 에이전트는 인간이 반복적인 GUI 작업을 수행하는 데 도움을 줄 수 있으며, 복잡한 작업을 실행 가능한 작업으로 분해하는 데 효과적인 작업 계획이 필수적입니다. 작은 오픈 소스 MLLM은 상업용 대규모 모델에 비해 비용 효율적이고 개인 정보 보호 측면에서 유리하지만, 계획 능력 부족과 제한적인 웹사이트 간 일반화 능력이 단점입니다. 이러한 한계를 해결하기 위해, 우리는 계획 경험 탐색 및 활용(PEEU) 방법을 제안합니다. 이 방법은 에이전트가 환경을 자율적으로 탐색하여 경험을 발견하고, 후향적 경험을 활용하여 엄격하게 정렬된 고수준 학습 데이터를 생성합니다. 이러한 성능의 원인이 되는 일반화 행동을 정량적으로 분석하기 위해, 우리는 작업 분해 계층 구조 분석 프레임워크(TDHAF)를 제안하여 세 가지 수준(낮음, 중간, 높음)의 작업에서 구성적 일반화를 체계적으로 연구합니다. 우리의 분석 결과, 낮은 수준의 기본적인 기술 습득이 반드시 고수준의 계획 능력으로 이어지지 않으며, 반대로 고수준 작업 훈련은 더 강력한 OOD (Out-of-Distribution) 일반화 능력을 제공한다는 것을 알 수 있습니다. 실제 벤치마크에서의 실험 결과는 PEEU의 우수한 효과를 입증합니다. 저희가 개발한 7B 모델은 30.6%의 정확도를 달성하여, 훨씬 더 큰 Qwen2.5-VL-32B 모델보다 뛰어난 성능을 보였습니다. 이러한 결과는 후향적 고수준 작업을 구성하고 경험을 활용하는 것이 작은 MLLM의 OOD 계획 능력에 매우 중요하다는 것을 보여줍니다.
Multimodal web agents can assist humans in operating repetitive GUI tasks, where effective task planning is essential for decomposing complex tasks into executable actions. While small open source MLLMs are cost efficient and privacy preserving compared with commercial large models, they suffer from weak planning and limited cross website generalization. To address these limitations, we introduce the planning experience exploration and utilization (PEEU) method, which autonomously explores environments to discover experiences and utilizes hindsight experience to synthesize strictly aligned, high level training data. To quantitatively analyze the generalization behaviors driving this performance, we propose the task decomposition hierarchical analysis framework (TDHAF) to systematically study compositional generalization across three task granularities: low, middle and high levels. Our analysis reveals that mastering low level atomic skills does not guarantee high level planning competence, while high level task training yields stronger OOD generalization. Experiments on real world benchmarks demonstrate PEEU's superior effectiveness: our 7B model achieves 30.6% accuracy, outperforming the much larger Qwen2.5-VL-32B model. These demonstrate constructing hindsight high level tasks and leveraging experiences is crucial for OOD planning abilities of small MLLMs.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.