2606.05806v1 Jun 04, 2026 cs.AI

도구가 실패할 때: LLM 에이전트의 동적 재계획 및 이상 복구 성능 측정

When Tools Fail: Benchmarking Dynamic Replanning and Anomaly Recovery in LLM Agents

Shuaiqiang Wang
Shuaiqiang Wang
Citations: 2,668
h-index: 21
Lingyong Yan
Lingyong Yan
Baidu Inc.
Citations: 1,503
h-index: 17
Xiang Li
Xiang Li
Citations: 19
h-index: 2
Yucheng Shen
Yucheng Shen
Citations: 47
h-index: 3
Dongsheng Zhu
Dongsheng Zhu
Citations: 94
h-index: 3
Dawei Yin
Dawei Yin
Citations: 106
h-index: 5
Xucheng Ma
Xucheng Ma
Citations: 0
h-index: 0
Yukun Zhao
Yukun Zhao
Citations: 96
h-index: 3

기존의 성능 측정 도구는 LLM에서 도구를 활용한 추론(TIR)을 이상적인 '정상 경로'를 기준으로 평가하며, 실제 환경에서 발생하는 도구 오류를 간과하는 경향이 있습니다. 본 연구에서는 TIR 에이전트의 동적 경로 탐색 및 오류 복구 성능을 평가하기 위한 벤치마크인 ToolMaze를 소개합니다. ToolMaze는 체계적인 재계획과 무분별한 시행착오를 구분하기 위해, 방향 비선 그래프(DAG) 기반의 위상 복잡성과 도구 변경 유형에 대한 $2 imes 2$ 분류 (명시적/암묵적, 일시적/영구적)라는 두 가지 차원을 사용합니다. 실험 결과는 대부분의 모델에서 오류가 성능 저하를 유발하며, 특히 암묵적인 의미 오류(implicit semantic failures) 하에서 성능 저하가 가장 심각한 것으로 나타났습니다. 시스템적인 신뢰 부족으로 인해, 손상된 출력에 대한 복구율(Perturbation Recovery Rate, PRR)이 이러한 상황에서 약 37% 감소했으며, 복잡한 위상 구조는 에이전트를 무의미한 시행착오 루프에 빠뜨립니다. 중요한 점은 에이전트의 오류 내성 기능이 기본 작업 실행보다 $3.66$배 더 느리게 개선되며, 이는 동적 재계획이 모델 크기 확장이나 프롬프트 최적화로 해결할 수 없는 별도의 병목 현상임을 시사합니다. 데이터 및 코드는 https://github.com/Zhudongsheng75/ToolMaze 에서 확인할 수 있습니다.

Original Abstract

Existing benchmarks evaluate Tool-Integrated Reasoning (TIR) in LLMs on idealized ''happy paths'', largely overlooking real-world tool failures. We introduce ToolMaze, a benchmark for dynamic path discovery and error recovery in TIR agents. To separate systematic replanning from blind trial-and-error, ToolMaze adopts a two-dimensional design: DAG-based topological complexity and a $2 \times 2$ taxonomy of tool perturbations (explicit/implicit, transient/permanent). Evaluations show that perturbations degrade performance across nearly all models, with the sharpest drops under implicit semantic failures. Driven by systemic over-trust in corrupted outputs, Perturbation Recovery Rate (PRR) plummets by around 37\% in these scenarios, while complex topologies trap agents in futile trial-and-error loops. Crucially, agentic fault-tolerance improves with model scale $3.66\times$ slower than basic task execution, highlighting dynamic replanning as a distinct bottleneck unaddressed by model scaling or prompting. Data and code are available at https://github.com/Zhudongsheng75/ToolMaze.

2 Citations
0 Influential
35.993061443341 Altmetric
11.0 Score
Original PDF
2

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!