Workflow-GYM: 실제 업무 환경에서 컴퓨터 사용 에이전트의 장기적인 작업 수행 능력 평가를 위한 연구
Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields
최근 몇 년 동안 인공지능 에이전트는 점점 더 복잡하고 실질적인 작업을 처리하는 방향으로 빠르게 발전해 왔습니다. 그러나 기존 벤치마크는 에이전트가 다양한 분야에서 장기적이고 가치 있는 전문 업무 워크플로우를 수행하기 위해 그래픽 사용자 인터페이스(GUI)를 얼마나 효과적으로 사용할 수 있는지 평가하지 못하는 경우가 많습니다. 현재 GUI 벤치마크는 여전히 주로 범용 소프트웨어, 비교적 간단한 애플리케이션 및 단기적인 작업에 초점을 맞추고 있으며, 현대 에이전트가 사용자의 지시를 따라 특정 분야의 전문 소프트웨어를 자율적으로 운영하고 종단 간(end-to-end) 방식으로 경제적 가치를 창출하는 작업을 수행할 수 있는지 여부는 아직 명확하게 밝혀지지 않았습니다. 이러한 격차를 해소하기 위해, 우리는 전문 분야 및 특수 소프트웨어 환경에 중점을 둔 장기적인 GUI 작업 평가를 위한 벤치마크인 Workflow-GYM을 소개합니다. 최첨단 모델에 대한 광범위한 실험을 통해, 가장 강력한 모델조차도 30% 미만의 낮은 성공률을 보이는 것을 확인했으며, 이는 전문 분야의 장기적인 GUI 워크플로우가 현재의 GUI 에이전트에게 여전히 매우 어려운 과제임을 시사합니다. 추가 분석 결과, 현재 에이전트는 장기적인 워크플로우 일관성을 유지하는 데 어려움을 겪으며, 종종 작업 단계 누락, 오류 전파, 목표 편향, 그리고 전문 소프트웨어 환경에 대한 이해 부족을 보이는 것으로 나타났습니다. 이러한 연구 결과는 현재 에이전트 시스템의 한계를 보여주며, 차세대 GUI-에이전트 연구를 위한 중요한 방향을 제시합니다.
Recent years have witnessed the rapid evolution of AI agents toward handling increasingly complex, real-world tasks. However, existing benchmarks rarely evaluate whether agents can operate graphical user interfaces to complete long-horizon, high-value professional workflows across diverse domains. Current GUI benchmarks still predominantly focus on general-purpose software, relatively simple applications, and short-horizon tasks, leaving it largely unknown whether modern agents can follow user instructions to autonomously operate domain-specific professional software and accomplish economically valuable work in an end-to-end manner. To bridge this gap, we introduce Workflow-GYM, a benchmark for long-horizon GUI tasks centered on professional domains and specialized software environments. Through extensive experiments on state-of-the-art models, we find that even the strongest models achieve only slightly above 30% success rates, highlighting that professional long-horizon GUI workflows remain highly challenging for current GUI agents. Further analysis reveals that current agents struggle to maintain long-horizon workflow consistency, frequently exhibiting workflow stage omission, error propagation, objective drift, and insufficient understanding of professional software environments. Our findings provide important insights into the limitations of current agent systems and suggest key directions for the next generation of GUI-agent research.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.