2605.29218v1 May 28, 2026 cs.AI

GTA: 웹 에이전트를 위한 확장 가능한 장기 작업 생성

GTA: Generating Long-Horizon Tasks for Web Agents at Scale

Prafulla Kumar Choubey
Prafulla Kumar Choubey
Texas A&M Univeristy
Citations: 931
h-index: 17
Kung-Hsiang Huang
Kung-Hsiang Huang
Citations: 204
h-index: 7
Chien-Sheng Wu
Chien-Sheng Wu
Citations: 40
h-index: 4
Tenghao Huang
Tenghao Huang
Citations: 1,621
h-index: 9
Yilun Zhou
Yilun Zhou
Citations: 165
h-index: 8
Muhao Chen
Muhao Chen
Citations: 37
h-index: 4
Jonathan May
Jonathan May
Citations: 37
h-index: 5

언어 모델과 브라우징 및 도구 사용 기능을 결합하는 웹 에이전트는 오픈 웹 어시스턴트로서의 잠재력을 보여줍니다. 그러나 현재 진행 상황은 주로 확장 가능한 프로세스 수준의 감독 부족으로 인해 제한됩니다. 기존 벤치마크는 대부분 수동으로 구성되며, 중간 경로 없이 대략적인 시작-목표 주석만 제공합니다. 최근 자동 생성 시도는 여전히 비용이 많이 들고 편향되어 있으며 표면적입니다. 이러한 한계는 현실적인 다단계, 페이지 간 작업을 일반화해야 하는 에이전트의 안정적인 학습 및 평가를 방해합니다. 우리는 크롤링, 검색 기반 시드, 컨텍스트 내 생성 및 자동 품질 관리를 통합하여 실제 작업과 실행 가능한 경로를 함께 생성하는 확장 가능한 프레임워크인 GTA를 소개합니다. 이 설계는 효율성을 높이기 위해 크롤링과 생성을 분리하고, 사이트 그래프에 작업을 기반하여 구성 가능성을 강화하며, 결정적인 반복 및 체계적인 검증을 통해 밀집된 감독을 보장합니다. 우리는 전자 상거래, 정부, 포럼 및 뉴스 웹사이트를 포함한 50개 이상의 웹사이트에서 다국어 및 다단계 지원으로 이 파이프라인을 구현했습니다. 결과적으로 생성된 벤치마크는 인간-에이전트 성능 간의 상당한 격차를 보여주며, 상세한 분석을 가능하게 합니다. 우리의 기여는 세 가지입니다: (i) 다단계 웹 에이전트 작업 생성을 공식화하고, (ii) 효율적이고 검증된 자동 데이터 생성 파이프라인을 제안하며, (iii) 재현 가능한 평가를 위한 동적 벤치마크를 공개합니다.

Original Abstract

Web agents, which couple language models with browsing and tool-use capabilities, show promise as open web assistants. Yet progress is increasingly limited by the lack of scalable, process-level supervision. Existing benchmarks are largely manually constructed, providing only coarse start-goal annotations without intermediate trajectories, while recent automatic generation efforts remain expensive, biased, and shallow. These limitations prevent reliable training and evaluation of agents that must generalize to realistic, multi-hop, cross-page tasks. We introduce a scalable framework, GTA, that integrates crawling, retrieval-based seeding, in-context generation, and automated quality control to produce realistic tasks paired with executable trajectories. This design decouples crawling from generation for greater efficiency, grounds tasks in the site graph to enforce compositionality, and ensures dense supervision through deterministic replays and systematic validation. We instantiate the pipeline on over 50 websites covering e-commerce, government, forums, and news, with multilingual and multi-hop coverage. The resulting benchmark reveals a significant human-agent performance gap and enables detailed diagnostics. Our contributions are three-fold: (i) formalizing multi-hop web-agent task generation, (ii) proposing an efficient and validated pipeline for automatic data creation, and (iii) releasing a dynamic benchmark with reproducible evaluation.

0 Citations
0 Influential
8.5 Altmetric
42.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!