2605.29225v1 May 28, 2026 cs.AI

BenchTrace: LLM 에이전트의 반사 능력 및 제어된 진화를 테스트하기 위한 벤치마크

BenchTrace: A Benchmark for Testing Reflection Ability and Controlled Evolution in LLM Agents

Fei Cheng
Fei Cheng
Graduate School of Informatics, Kyoto University
Citations: 909
h-index: 15
Jiahao Huang
Jiahao Huang
Citations: 2
h-index: 1
Junfeng Jiang
Junfeng Jiang
Citations: 122
h-index: 5
Akiko Aizawa
Akiko Aizawa
Citations: 112
h-index: 5
Zefan Yu
Zefan Yu
Citations: 65
h-index: 2

자기 진화 에이전트는 과거 실패를 되돌아보며 시간이 지남에 따라 개선되지만, 기존 평가 방법은 다음과 같은 두 가지 측면에서 한계가 있습니다. 첫째, 작업 수행 점수만을 측정하여 반사 능력의 품질을 알 수 없으며, 둘째, 에이전트 자체의 실행 결과를 기반으로 하여 특정 실패 패턴을 목표로 하는 메커니즘을 제공하지 않습니다. 본 논문에서는 LLM 에이전트의 자기 진화 능력을 평가하기 위한 벤치마크인 extbf{BenchTrace}를 제안합니다. BenchTrace는 여섯 가지 다양한 작업에 걸쳐 1,821개의 주석 처리된 에피소드로 구성된 스냅샷-반사 데이터셋을 기반으로 하며, extbf{반사 평가(Reflection Evaluation)}는 목표 QA 작업을 통해 실패 식별 능력을 측정하고, extbf{진화 평가(Evolution Evaluation)}는 과거의 실패 경험이 제어된 자기 진화 시뮬레이션에서 회피 행동으로 이어지는지 테스트합니다. BenchTrace를 기반으로, 대상 실패 사례를 성공적으로 회피하는 테스트 케이스의 비율을 측정하는 새로운 평가 지표인 extbf{실패 회피율 (failure avoidance rate, FAR)}을 제안합니다. Qwen3-32B와 GPT-4.1 모델에 대한 실험 결과, 두 모델 모두 반사 평가에서 30% 미만의 전체 성공률을 보였으며, 특히 진단 능력이 주요 병목 현상으로 작용했습니다. 진화 평가는 자기 진화 방법이 일반적으로 비진화 기준보다 FAR을 향상시키지만, 잡음 에피소드가 누적됨에 따라 에이전트가 초기 학습 내용을 잊어버리고, 특정 맥락 너머의 반사를 일반화하지 못하여 작업 맥락 간에 부정적인 영향을 미치는 것을 보여줍니다. 또한 상관 분석 결과, 완전히 정확한 반사가 높은 FAR과 강하게 관련되어 있음을 확인했습니다. BenchTrace는 현재 자기 진화 접근 방식의 구체적인 한계를 드러내며, 대상 평가를 위한 제어되고 모델 독립적인 프레임워크를 제공합니다.

Original Abstract

Self-evolving agents improve over time by reflecting on past failures, but existing evaluation is limited in two ways: it measures only task scores, leaving reflection quality unknown, and it relies on agents' own episode runs, offering no mechanism to target specific failure patterns. We present \textbf{BenchTrace}, a benchmark for evaluating self-evolution ability in LLM agents. BenchTrace is built on a snapshot-reflection dataset of 1,821 annotated episodes spanning six diverse tasks, and comprises a \textbf{Reflection Evaluation} that probes failure identification through targeted QA tasks, and an \textbf{Evolution Evaluation} that tests whether past failure experience translates into avoidance behavior in a controlled self-evolution simulation. Building on BenchTrace, we propose \textbf{failure avoidance rate (FAR)}, a new evaluation metric measuring the fraction of test cases in which the agent successfully avoids the target failure instance. Experiments with Qwen3-32B and GPT-4.1 reveal that both models fall below a 30\% end-to-end pass rate on reflection evaluation, with diagnosis as the primary bottleneck. Evolution evaluation shows that self-evolution methods generally improve FAR over the non-evolving baseline, but agents forget early lessons as noise episodes accumulate, and agents fail to generalize their reflections beyond the specific context, causing negative transfer across task contexts. Our correlation analysis further reveals that only a fully correct reflection is strongly associated with higher FAR. BenchTrace exposes concrete limits of current self-evolution approaches and provides a controlled, model-agnostic framework for targeted evaluation.

1 Citations
0 Influential
7.5 Altmetric
38.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!