2604.17849v1 Apr 20, 2026 cs.AI

컴퓨터 사용 에이전트의 신뢰성에 대한 연구

On the Reliability of Computer Use Agents

Gonzalo Gonzalez-Pumariega
Gonzalo Gonzalez-Pumariega
Citations: 228
h-index: 9
Saaket Agashe
Saaket Agashe
Citations: 456
h-index: 7
Jiachen Yang
Jiachen Yang
Citations: 329
h-index: 4
Ang Li
Ang Li
Citations: 60
h-index: 3
X. Wang
X. Wang
Citations: 361
h-index: 6

컴퓨터 사용 에이전트는 웹 탐색, 데스크톱 자동화 및 소프트웨어 상호 작용과 같은 실제 작업에서 빠르게 발전하여, 경우에 따라 인간의 성능을 능가하기도 합니다. 그러나 작업과 모델이 동일하더라도, 한 번 성공한 에이전트가 동일한 작업을 반복 실행할 때 실패할 수 있습니다. 이는 근본적인 질문을 제기합니다. 에이전트가 한 번 작업을 성공적으로 수행했다면, 무엇이 그 에이전트가 해당 작업을 안정적으로 수행하는 것을 방해하는가? 본 연구에서는 컴퓨터 사용 에이전트의 신뢰성 문제를 실행 중 발생하는 확률적 요소, 작업 정의의 모호성, 그리고 에이전트 행동의 변동성이라는 세 가지 요소를 통해 분석합니다. 우리는 OSWorld 환경에서 동일한 작업을 반복 실행하고, 설정 간의 작업 수준 변화를 파악하기 위한 통계적 검정을 병행하여 이러한 요인들을 분석했습니다. 분석 결과, 신뢰성은 작업 정의 방식과 에이전트 행동의 실행 간 변동성에 모두 의존한다는 것을 확인했습니다. 이러한 결과는 에이전트를 반복 실행하여 평가하고, 에이전트가 상호 작용을 통해 작업의 모호성을 해결할 수 있도록 하며, 실행 간 안정성이 높은 전략을 선호해야 할 필요성을 시사합니다.

Original Abstract

Computer-use agents have rapidly improved on real-world tasks such as web navigation, desktop automation, and software interaction, in some cases surpassing human performance. Yet even when the task and model are unchanged, an agent that succeeds once may fail on a repeated execution of the same task. This raises a fundamental question: if an agent can succeed at a task once, what prevents it from doing so reliably? In this work, we study the sources of unreliability in computer-use agents through three factors: stochasticity during execution, ambiguity in task specification, and variability in agent behavior. We analyze these factors on OSWorld using repeated executions of the same task together with paired statistical tests that capture task-level changes across settings. Our analysis shows that reliability depends on both how tasks are specified and how agent behavior varies across executions. These findings suggest the need to evaluate agents under repeated execution, to allow agents to resolve task ambiguity through interaction, and to favor strategies that remain stable across runs.

2 Citations
0 Influential
4.5 Altmetric
24.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!