컴퓨터 사용 에이전트의 신뢰성에 대한 연구
On the Reliability of Computer Use Agents
컴퓨터 사용 에이전트는 웹 탐색, 데스크톱 자동화 및 소프트웨어 상호 작용과 같은 실제 작업에서 빠르게 발전하여, 경우에 따라 인간의 성능을 능가하기도 합니다. 그러나 작업과 모델이 동일하더라도, 한 번 성공한 에이전트가 동일한 작업을 반복 실행할 때 실패할 수 있습니다. 이는 근본적인 질문을 제기합니다. 에이전트가 한 번 작업을 성공적으로 수행했다면, 무엇이 그 에이전트가 해당 작업을 안정적으로 수행하는 것을 방해하는가? 본 연구에서는 컴퓨터 사용 에이전트의 신뢰성 문제를 실행 중 발생하는 확률적 요소, 작업 정의의 모호성, 그리고 에이전트 행동의 변동성이라는 세 가지 요소를 통해 분석합니다. 우리는 OSWorld 환경에서 동일한 작업을 반복 실행하고, 설정 간의 작업 수준 변화를 파악하기 위한 통계적 검정을 병행하여 이러한 요인들을 분석했습니다. 분석 결과, 신뢰성은 작업 정의 방식과 에이전트 행동의 실행 간 변동성에 모두 의존한다는 것을 확인했습니다. 이러한 결과는 에이전트를 반복 실행하여 평가하고, 에이전트가 상호 작용을 통해 작업의 모호성을 해결할 수 있도록 하며, 실행 간 안정성이 높은 전략을 선호해야 할 필요성을 시사합니다.
Computer-use agents have rapidly improved on real-world tasks such as web navigation, desktop automation, and software interaction, in some cases surpassing human performance. Yet even when the task and model are unchanged, an agent that succeeds once may fail on a repeated execution of the same task. This raises a fundamental question: if an agent can succeed at a task once, what prevents it from doing so reliably? In this work, we study the sources of unreliability in computer-use agents through three factors: stochasticity during execution, ambiguity in task specification, and variability in agent behavior. We analyze these factors on OSWorld using repeated executions of the same task together with paired statistical tests that capture task-level changes across settings. Our analysis shows that reliability depends on both how tasks are specified and how agent behavior varies across executions. These findings suggest the need to evaluate agents under repeated execution, to allow agents to resolve task ambiguity through interaction, and to favor strategies that remain stable across runs.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.