2607.28367v1 Jul 30, 2026 cs.AI

컴퓨터 사용 에이전트의 성능 평가 지표가 잘못된 점수 결과를 초래하는 방식

How Benchmarks Mis-Score Computer-Use Agents

Zihan Dong
Zihan Dong
Citations: 107
h-index: 5
Zhiyuan Ma
Zhiyuan Ma
Citations: 19
h-index: 2
Rui Qian
Rui Qian
Citations: 643
h-index: 9
Zekun Wang
Zekun Wang
Citations: 26
h-index: 2
Ruifang Qian
Ruifang Qian
Citations: 19
h-index: 1
Qi Zhan
Qi Zhan
Citations: 3
h-index: 1
Yunqing Li
Yunqing Li
Citations: 6
h-index: 1
Zirou Liu
Zirou Liu
Citations: 0
h-index: 0
Ruixuan Deng
Ruixuan Deng
Citations: 92
h-index: 3

웹 브라우징 및 데스크톱 소프트웨어 운영을 위해 컴퓨터 사용 에이전트(CUA)가 널리 활용되고 있지만, 이러한 에이전트들의 성능 평가는 여전히 불안정한 스크립트 기반 평가 시스템에 의해 주로 이루어지고 있습니다. 이러한 평가는 과제가 오래되었거나, 경로가 중요한 시각적 증거를 누락하거나, 평가자가 유효한 대안을 거부하거나, 종합 보고서가 실패 원인을 숨길 수 있는 파이프라인의 결과물입니다. 본 연구에서는 이러한 문제들을 작업 구성, 경로 관찰, 점수 산정 및 보고라는 신뢰성 프레임워크로 분류합니다. 이후, 5개의 웹, 기업 워크플로우 및 데스크톱 제어 벤치마크에서 수집된 150건의 실패 사례 평가 경로를 분석한 결과, FAIL 판정의 15.3%가 잘못된 것으로 나타났습니다. 그 중 10.7%는 평가자의 오판(false negative)이었고, 4.7%는 작업 자체에 오류가 있었습니다. 실제 실패 사례에 대한 세 단계 진단 분류법 분석 결과, 검증/피드백 실패 및 계획 실패가 실행/접지 오류보다 더 흔하게 발생하며, 단일 성공률 지표로는 이러한 현상을 설명하기 어렵습니다. 본 연구의 결과를 바탕으로 새로운 장기적 CUA 벤치마크에 대한 시사점을 도출하고, CUA 평가를 위한 단계별 설계 규칙을 제시합니다.

Original Abstract

Computer-use agents (CUA) are being deployed to browse the web and operate desktop software, yet their benchmark scores are still commonly produced by brittle scripted oracles. A score is the output of a pipeline in which tasks can be stale, trajectories can omit decisive visual evidence, evaluators can reject valid alternatives, and aggregate reports can hide the cause of failure. We organize these problems into a reliability framework spanning task construction, trajectory observation, scoring, and reporting. We then audit 150 public failure-scored trajectories from five web, enterprise-workflow, and desktop-control benchmarks, find that 15.3\% of FAIL verdicts are wrong: 10.7\% are evaluator false negatives and 4.7\% are broken tasks. For genuine failures, a three-tier diagnostic taxonomy shows that verification/feedback and planning failures dominate execution/grounding errors, while a single scalar success rate can not explain. We connect these findings to newer long-horizon CUA benchmarks and derive stage-specific design rules for CUA evaluation.

0 Citations
0 Influential
4.5 Altmetric
22.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!