인공지능 에이전트의 신뢰성에 대한 연구
Towards a Science of AI Agent Reliability
인공지능 에이전트는 중요한 작업을 수행하기 위해 점점 더 많이 활용되고 있습니다. 표준 벤치마크에서 정확도가 향상되는 것처럼 보이는 것은 빠른 발전을 나타내지만, 많은 에이전트가 실제 환경에서 여전히 실패하는 경우가 많습니다. 이러한 불일치는 현재 평가 방식의 근본적인 한계를 보여줍니다. 에이전트의 행동을 단일 성공 지표로 압축하는 것은 중요한 운영상의 결함을 가립니다. 특히, 에이전트가 실행 과정에서 일관성을 유지하는지, 외부 요동에 얼마나 잘 대응하는지, 예상 가능한 방식으로 실패하는지, 그리고 오류의 심각성이 제한되어 있는지 등을 고려하지 않습니다. 안전이 중요한 공학 분야의 원칙에 기반하여, 우리는 일관성, 강건성, 예측 가능성, 안전성이라는 네 가지 핵심 차원을 따라 에이전트의 신뢰성을 분석하는 열두 가지 구체적인 지표를 제안하여 종합적인 성능 프로필을 제공합니다. 두 가지 상호 보완적인 벤치마크에서 14개의 모델을 평가한 결과, 최근의 성능 향상이 신뢰성 측면에서 미미한 개선만을 가져왔음을 확인했습니다. 이러한 지속적인 한계를 드러냄으로써, 우리의 지표는 기존 평가 방식에 보완적인 역할을 하며, 에이전트가 어떻게 작동하고, 성능이 저하되며, 실패하는지에 대한 이해를 돕는 도구를 제공합니다.
AI agents are increasingly deployed to execute important tasks. While rising accuracy scores on standard benchmarks suggest rapid progress, many agents still continue to fail in practice. This discrepancy highlights a fundamental limitation of current evaluations: compressing agent behavior into a single success metric obscures critical operational flaws. Notably, it ignores whether agents behave consistently across runs, withstand perturbations, fail predictably, or have bounded error severity. Grounded in safety-critical engineering, we provide a holistic performance profile by proposing twelve concrete metrics that decompose agent reliability along four key dimensions: consistency, robustness, predictability, and safety. Evaluating 14 models across two complementary benchmarks, we find that recent capability gains have only yielded small improvements in reliability. By exposing these persistent limitations, our metrics complement traditional evaluations while offering tools for reasoning about how agents perform, degrade, and fail.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.