2608.06346v1 Aug 06, 2026 cs.AI

TRAJDEBUG: 에이전트의 장기 실행 경로에서 치명적인 오류를 식별하기 위한 오류 수명 주기 추적

TRAJDEBUG: Tracing Error Lifecycle to Identify Critical Failures in Long-Horizon Agent Trajectories

Hao Peng
Hao Peng
Citations: 540
h-index: 10
Yunjia Qi
Yunjia Qi
Citations: 336
h-index: 8
Lei Hou
Lei Hou
Citations: 452
h-index: 8
Juanzi Li
Juanzi Li
Citations: 932
h-index: 13
Yixian Liu
Yixian Liu
Citations: 4
h-index: 1
Zhichao Hu
Zhichao Hu
Citations: 50
h-index: 5
Yuhong Liu
Yuhong Liu
Citations: 318
h-index: 4
Xiaozhi Wang
Xiaozhi Wang
Citations: 108
h-index: 4
Bin Xu
Bin Xu
Citations: 135
h-index: 4
Richeng Xuan
Richeng Xuan
Citations: 108
h-index: 4
Zehua Yin
Zehua Yin
Citations: 0
h-index: 0
Xin Shi
Xin Shi
Citations: 361
h-index: 2
Songyuanyi Lu
Songyuanyi Lu
Citations: 0
h-index: 0

LLM 기반 에이전트 시스템은 복잡한 영역에서 놀라운 능력을 보여주지만, 연쇄적인 오류와 디버깅의 어려움을 겪습니다. 치명적인 오류 탐지는 실패한 실행 경로에서 최종 실패를 유발하는 초기 오류 단계를 찾는 것을 목표로 합니다. 그러나 이 분야는 두 가지 주요 과제에 직면합니다. 첫째, 긴 실행 경로는 개별 오류를 식별하기 어렵게 만듭니다. 왜냐하면 특정 단계의 오류 여부를 판단할 증거가 멀리 떨어진 명령어, 관찰 및 이전 컨텍스트에 흩어져 있을 수 있기 때문입니다. 둘째, 실패한 실행 경로에는 종종 다양한 결과로 이어지는 여러 개의 국소적인 오류가 포함되어 있으며, 이 중에서 일부만이 최종 실패를 유발하는 원인이 됩니다. 본 연구에서는 TrajDebug라는 오류 수명 주기 추적 프레임워크를 제안합니다. TrajDebug는 다중 수준의 히스토리 압축과 증거 기반 오류 식별을 통해 장기 실행 경로에서 발생하는 오류를 효율적으로 탐지하며, 각 오류의 해결 상태 및 최종적인 영향까지 추적하여 원인 규명에 도움을 줍니다. 또한, Tau2Bench 및 SWE-Bench Pro에서 수동으로 주석이 달린 486개의 실패한 실행 경로로 구성된 TrajErrBench 벤치마크를 구축했습니다. 다양한 에이전트 벤치마크에서의 실험 결과, TrajDebug는 기존 방법보다 뛰어난 성능을 보였으며, 응용 연구를 통해 진단 결과가 하위 에이전트의 성공률 향상을 위한 실질적인 피드백을 제공한다는 것을 확인했습니다. 본 연구에서 개발된 코드와 데이터는 추가 연구를 위해 공개될 예정입니다.

Original Abstract

LLM-based agentic systems have shown remarkable capabilities in complex domains, while suffering from cascading errors and difficulty in debugging. Critical error detection aims to locate the earliest error step in a failed trajectory that is responsible for the final failure. However, progress faces two main challenges. First, long trajectories make it difficult to identify individual errors, since the evidence for judging a step may be scattered across distant instructions, observations, and prior context. Second, failed trajectories often contain multiple local errors with different downstream effects, only some of which remain responsible for the final failure. In this work, we propose TrajDebug, an error-lifecycle tracing framework that addresses long-trajectory error discovery with multi-granularity history compression and evidence-based error identification, and supports critical attribution by tracing each error's resolution status and terminal impact. We further construct TrajErrBench, a benchmark of 486 manually annotated failed trajectories from Tau2Bench and SWE-Bench Pro, covering realistic tool-use and coding scenarios. Experiments across diverse agent benchmarks show that TrajDebug achieves the best overall performance over existing baselines, and application studies further demonstrate that its diagnoses provide actionable feedback for improving downstream agent success. We will release the codes and data to facilitate further research.

0 Citations
0 Influential
6.5 Altmetric
32.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!