WeaveBench: 하이브리드 인터페이스를 사용하는 컴퓨터 활용 에이전트를 위한 장기적인 실세계 기반 평가 도구
WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces
컴퓨터 활용 에이전트(CUA)는 시각적 데스크톱 제어, 명령줄 실행, 코드 편집, 브라우저 및 외부 도구를 결합한 환경에서 작동하는 경우가 많습니다. 그러나 기존의 벤치마크는 이러한 인터페이스를 독립적인 기능으로 평가하는 경향이 있어, 장기적인 관점에서 여러 인터페이스를 조율하는 능력은 충분히 검증되지 않았습니다. 이에 따라, 우리는 실제 사용자 요청과 공개적으로 검증 가능한 결과물을 기반으로 한 114개의 작업과 8가지 실세계 업무 영역을 포함하는 장기적인 하이브리드 인터페이스 벤치마크인 WeaveBench를 소개합니다. 각 작업은 에이전트가 단일 실행 경로 내에서 GUI 관찰/작업과 CLI/코드 작업을 결합해야 합니다. 우리는 최소한의 데스크톱 제어 플러그인이 추가된 실제 Ubuntu 데스크톱 환경에서 배포된 CLI-에이전트 런타임에서 이러한 작업을 평가했습니다. 또한, 결과물, 파일, 스크린샷, 로그 및 작업 추적 정보를 검사하고 조작된 시각 자료 또는 하드 코딩된 지표와 같은 비정상적인 동작을 감지하는 trajectory-aware 심사 시스템을 제안합니다. 최첨단 모델-런타임 조합의 경우, 최고 PassRate는 41.2%에 불과하며, 이는 벤치마크가 아직 충분히 활용되지 않았음을 보여줍니다. trajectory-aware 심사 시스템은 또한 결과 중심 평가가 에이전트 성능을 과대평가한다는 것을 추가적으로 밝혀냅니다. 전반적으로, WeaveBench는 CUA 평가의 중요한 격차를 드러내며, 에이전트가 다양한 GUI, CLI 및 코드 작업을 장기적인 실세계 작업에 걸쳐 조율할 수 있는지 측정하기 위한 효과적인 테스트 환경을 제공합니다.
Computer-use agents (CUAs) increasingly operate in runtimes that combine visual desktop control, command-line execution, code editing, browsers, and external tools. Existing benchmarks, however, often evaluate these interfaces as separable capabilities, leaving long-horizon cross-interface orchestration under-tested. Thus, we introduce WeaveBench, a long-horizon hybrid-interface benchmark with 114 tasks across 8 real-world work domains, grounded in real user requests and publicly verifiable artifacts. Each task requires agents to combine GUI observations/actions with CLI/code operations within a single trajectory. We evaluate these tasks on a real Ubuntu desktop inside deployed CLI-agent runtimes, augmented with a minimal desktop-control plugin. We also propose a companion trajectory-aware judge that inspects deliverables, files, screenshots, logs, and action traces, while detecting shortcut behaviors such as fabricated visual evidence or hard-coded metrics. Across frontier model-runtime pairings, the best PassRate reaches only 41.2%, showing the benchmark remains far from saturated. The trajectory-aware judge further reveals that outcome-only grading substantially overestimates agent performance. Overall, WeaveBench exposes a critical gap in CUA evaluation and provides an effective testbed to measure whether agents can orchestrate GUI, CLI, and code operations across long-horizon real-world tasks.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.