2607.28802v1 Jul 30, 2026 cs.AI

모델인가, 하드웨어인가? 에이전트 오류를 지역화하기 위한 상호작용 중심 분류 체계

Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures

Harsh Raj
Harsh Raj
Citations: 103
h-index: 4
Vipul Gupta
Vipul Gupta
Pennsylvania State University
Citations: 511
h-index: 14
Anas Mahmoud
Anas Mahmoud
Citations: 548
h-index: 4
Razvan-Gabriel Dumitru
Razvan-Gabriel Dumitru
Citations: 210
h-index: 5
Darvin Yi
Darvin Yi
Citations: 20
h-index: 3
Aakash Sabharwal
Aakash Sabharwal
Citations: 8
h-index: 1
Yunzhong He
Yunzhong He
Citations: 109
h-index: 6

기존의 평가는 종종 에이전트 오류를 시스템 수준의 결과로 간주하여, 오류가 발생한 원점과 어떤 개입이 에이전트 시스템을 개선할 수 있는지 흐릿하게 만듭니다. 이는 문제 해결 과제를 야기합니다. 동일한 가시적인 오류는 모델 재학습, 하드웨어 설계 변경, 환경 재설계 또는 벤치마크 수정 등 다양한 대응 방안을 필요로 할 수 있습니다. 에이전트의 행동은 모델, 하드웨어, 사용자, 도구, 메모리 및 환경 간의 상호작용에서 비롯되기 때문에, 결과 수준의 레이블만으로는 개선에 충분하지 않습니다. 대부분의 오류 분류 체계는 벤치마크에 특화되어 있고 공유 구조가 부족하여 이러한 문제를 해결하는 데 도움이 되지 않습니다. 본 연구에서는 오류를 발생시킨 상호작용을 기준으로 에이전트 오류를 지역화하고 책임 있는 구성 요소를 식별하는 상호작용 중심의 분류 체계를 제안합니다. 이 체계는 41가지 오류 모드를 두 구성 요소 간의 연결(edge)과 수리 위치를 나타내는 '오류 측면'에 따라 분류하여 조직했습니다. 이를 통해, 모델 관련 오류는 재학습 대상이 되도록 하고, 하드웨어 관련 오류는 스캐폴딩 및 도구 통합 수정으로 이어지며, 환경 또는 평가 관련 오류는 재설계를 요구하는 평가 조건임을 명확히 합니다. 이 체계는 코딩 지원 에이전트부터 장기적인 개인 비서 및 다중 에이전트 시스템에 이르기까지 다양한 에이전트 아키텍처에 적용 가능합니다. 제안하는 분류 체계는 공개 벤치마크, 모델 시스템 카드, 발표된 보고서 및 기록된 에이전트 동작 로그에서 얻은 예시를 통해 구체화되었으며, 독립적인 추론 에이전트를 평가자로 사용하여 재현성을 평가했습니다. 네 가지 최첨단 모델에 대해 가장 성능이 좋은 평가자는 인간이 부여한 범주 레이블과 Cohen's $κ=0.76$의 높은 일치도를 보였습니다. 이는 제안된 범주가 개인별 주석 작성자의 선호도보다는 공유 구조를 반영한다는 것을 시사합니다.

Original Abstract

Existing evaluations often reduce agent failures to system-level outcomes, obscuring where the fault originated and which intervention would improve the agent system. This creates a repair-assignment problem: the same visible failure may call for model post-training, harness engineering, environment redesign, or benchmark repair depending on its source. Because agent behavior emerges from interactions among models, harnesses, users, tools, memory, and environments, outcome-level labels are often insufficient for improvement. Most failure taxonomies do little to resolve this problem because they are benchmark-specific and lack a shared structure. We introduce an interaction-centric taxonomy that localizes failures to the interactions in which they originate and identifies the responsible component. It organizes 41 failure modes by assigning each to an edge between two components and a fault side indicating where the repair belongs. This makes the taxonomy actionable: model-side failures identify targets for post-training, harness-side failures point to scaffolding and tool-integration fixes, and environment or grader failures reveal evaluation conditions requiring redesign. The schema applies across agent architectures, from coding assistants to long-horizon personal assistants and multi-agent systems. We ground the taxonomy in worked examples from public benchmarks, model system cards, published reports, and logged agent trajectories, and evaluate its reproducibility using independent reasoning agents as judges. Across four frontier models, the strongest judge reaches Cohen's $κ=0.76$ against human category labels, suggesting that the categories capture shared structure rather than annotator-specific preferences.

0 Citations
0 Influential
7 Altmetric
35.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!