2607.18754v1 Jul 21, 2026 cs.AI

AgentDebugX: LLM 에이전트의 오류 감지, 원인 분석 및 복구를 위한 오픈 소스 도구

AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM Agents

Bingxuan Li
Bingxuan Li
Citations: 86
h-index: 3
Zhiguang Han
Zhiguang Han
Citations: 120
h-index: 4
Jiaxuan You
Jiaxuan You
Citations: 368
h-index: 6
Heng Ji
Heng Ji
Citations: 997
h-index: 4
Xuyan Ye
Xuyan Ye
Citations: 13
h-index: 2
Weijia Zhang
Weijia Zhang
Citations: 64
h-index: 1
James Zou
James Zou
Citations: 516
h-index: 9
Pan Lu
Pan Lu
Citations: 86
h-index: 3
Kunlun Zhu
Kunlun Zhu
Citations: 468
h-index: 7
Yuchen Zhao
Yuchen Zhao
Citations: 0
h-index: 0
Muxin Tian
Muxin Tian
Citations: 70
h-index: 2
Xiangru Tang
Xiangru Tang
Citations: 11,502
h-index: 30

LLM (Large Language Model) 에이전트의 오류는 오작동이 발생하는 시점이 실제 원인이 발생한 시점과 일치하지 않는 경우가 많아 디버깅하기 어렵습니다. 기존의 관측성 도구들은 실행 추적을 재현하지만, 근본 원인을 파악하거나 진단을 복구 과정으로 전환하는 데 큰 도움이 되지 못합니다. 본 논문에서는 Detect, Attribute, Recover, Rerun의 순환 과정을 통해 디버깅을 체계화하는 오픈 소스 프레임워크인 AgentDebugX를 소개합니다. 핵심 기술인 DeepDebug는 글로벌 경로 이해, 구조 기반 조사 및 상호 검증을 통해 다단계 근본 원인 진단을 수행합니다. Who and When 벤치마크에서 DeepDebug는 평가된 방법 중 가장 높은 정확도를 달성했습니다. 특히 qwen3.5-9b 모델에서는 28.8%의 정확도로 에이전트와 단계를 정확하게 식별하는 반면, 최상의 단일 패스 기준 모델은 21.7%에 그쳤습니다. GAIA 데이터셋에서 DeepDebug는 단일 실행으로 73개의 실패한 작업 중 13개를 복구했습니다. 이는 세 가지 독립적인 자체 수정 기준 모델 (4~6개 복구)보다 우수한 성능입니다. 결과적으로 전체 정확도가 55.8%에서 63.6%로 향상되었습니다. AgentDebugX는 파이썬 라이브러리, CLI (Command Line Interface), 웹 콘솔 및 설치 가능한 에이전트 스킬을 통해 위 워크플로우를 제공하며, 익명화된 오류 진단 및 복구 정보를 공유하고 재사용할 수 있는 Error Hub 기능도 제공합니다.

Original Abstract

LLM agent failures are difficult to debug because the step where an error surfaces is often not the one that caused it. Existing observability tools replay execution traces but provide little support for identifying the root cause or translating diagnosis into recovery. We present AgentDebugX, an open-source debugging framework that organizes debugging as a closed loop of Detect, Attribute, Recover, and Rerun. At its core, DeepDebug performs multi-turn root-cause diagnosis through global trajectory understanding, structure-guided investigation, and cross-examination. On the Who and When benchmark, DeepDebug achieves the best strict attribution accuracy among the evaluated methods on both tested open-weight backbones, reaching 28.8 percent exact agent-and-step accuracy on qwen3.5-9b versus 21.7 percent for the strongest single-pass baseline. On GAIA, DeepDebug repairs 13 of 73 failed tasks in a single rerun, compared with 4 to 6 for three decoupled self-correction baselines, improving overall accuracy from 55.8 percent to 63.6 percent. AgentDebugX exposes this workflow through a Python library, CLI, web console, and installable agentic skill, and provides an opt-in Error Hub for sharing scrubbed failure-diagnosis-repair bundles and reusing them as debugging memory.

1 Citations
0 Influential
15 Altmetric
76.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!