2606.24626v1 Jun 23, 2026 cs.AI

SAFARI: 능동적 조사를 통한 장기 호라이즌 에이전트 오류 원인 분석의 확장

SAFARI: Scaling Long Horizon Agentic Fault Attribution via Active Investigation

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Shi-Xiong Zhang
Shi-Xiong Zhang
Citations: 122
h-index: 5
Sambit Sahu
Sambit Sahu
Citations: 87
h-index: 5
Sangwoo Cho
Sangwoo Cho
Tencent US
Citations: 769
h-index: 13
Kushal Chawla
Kushal Chawla
Citations: 8
h-index: 2
Chenyang Zhu
Chenyang Zhu
Citations: 28
h-index: 3
Pengshan Cai
Pengshan Cai
University of Massachusetts, Amherst
Citations: 432
h-index: 7
Erin Babinsky
Erin Babinsky
Citations: 11
h-index: 2
Jiayu Yao
Jiayu Yao
Citations: 158
h-index: 4
Youbing Yin
Youbing Yin
Citations: 10
h-index: 2
N. Wolfe
N. Wolfe
Citations: 82
h-index: 4
Jingyu Wu
Jingyu Wu
Citations: 12
h-index: 2
Daben Liu
Daben Liu
Citations: 57
h-index: 4

자율 에이전트가 점점 더 복잡하고 다단계, 다중 에이전트 작업을 수행함에 따라, 실행 경로는 가장 큰 컨텍스트 윈도우의 제약을 뛰어넘게 되었습니다. 현재 에이전트 오류를 효과적으로 진단하는 방법은 전체 경로를 LLM의 컨텍스트 윈도우에 로드하는 방식으로, 이는 어텐션 희석 문제를 야기하며, 불가피하게 에이전트 추적이 컨텍스트 제한을 초과할 때 실패합니다. 이러한 문제를 해결하기 위해, 우리는 SAFARI(Scaling long-horizon Agentic Fault AttRibution via active Investigation)라는 프레임워크를 제안합니다. SAFARI는 선형적인 컨텍스트 로딩 방식을 도구 기반 진단 루프로 대체하며, LLM에 경로 세그먼트를 읽고 검색할 수 있는 특수 도구를 제공하고, 턴 간 추론을 위한 영구적인 단기 기억(STM)을 활용하여, 진단 정확도를 아키텍처의 컨텍스트 제한과 분리합니다. 실험 결과, SAFARI는 1백만 토큰 예산 내에서 Who&When 데이터셋에서 최첨단 기술보다 20% 더 우수한 성능을 보이며, 25천 토큰 예산 내에서 TRAIL GAIA 서브셋에서도 19% 더 높은 성능을 보여줍니다. 특히, SAFARI는 모델의 기본 컨텍스트 윈도우 너머 5배 떨어진 곳에 목표 오류가 존재하더라도 0.58의 정밀도를 유지하며, 이는 기존 평가 방법으로는 완전히 실패하는 경우입니다.

Original Abstract

As autonomous agents tackle increasingly complex multi-step, multi-agent tasks, their execution trajectories have scaled beyond the constraints of even the largest context windows. Current methods for effectively diagnosing agent failures load the full trajectory into an LLM's context window, which suffers from attention dilution and fails when agentic traces inevitably exceed context limits. To address this, we introduce SAFARI (Scaling long-horizon Agentic Fault AttRibution via active Investigation), a framework that replaces linear context loading with a tool-augmented diagnostic loop. By equipping LLMs with a specialized toolbox to read and search trajectory segments alongside a persistent Short-Term Memory (STM) for cross-turn reasoning, SAFARI effectively decouples diagnostic accuracy from architectural context limits. Our experiments demonstrate that SAFARI outperforms state-of-the-art results by 20% on the Who&When dataset within a 1M token budget, and by 19% on TRAIL GAIA subset on a 25K token budget. Most significantly, SAFARI maintains a 0.58 precision even when the target fault resides 5x beyond the model's native context window, a scenario where traditional evaluators fail entirely.

0 Citations
0 Influential
6.5 Altmetric
32.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!