2607.05202v1 Jul 06, 2026 cs.AI

EvoAgentBench: 능력 전송을 통한 에이전트 자기 진화 성능 평가

EvoAgentBench: Benchmarking Agent Self-Evolution via Ability Transfer

Chuanrui Hu
Chuanrui Hu
Citations: 32
h-index: 3
Xingze Gao
Xingze Gao
Citations: 27
h-index: 2
Yi Bai
Yi Bai
Citations: 29
h-index: 2
Yafeng Deng
Yafeng Deng
Citations: 35
h-index: 3
Hongda Chen
Hongda Chen
Citations: 35
h-index: 3
Yunyun Han
Yunyun Han
Citations: 11
h-index: 2
Pengfei Yao
Pengfei Yao
Citations: 0
h-index: 0
Zhao Wang
Zhao Wang
Citations: 0
h-index: 0
Zhengwei Wu
Zhengwei Wu
Citations: 0
h-index: 0
Xiaofeng Cong
Xiaofeng Cong
Citations: 572
h-index: 14
Jie Gui
Jie Gui
Citations: 0
h-index: 0
Teng Li
Teng Li
Citations: 18
h-index: 2

장기적인 LLM 시스템에서 에이전트의 자기 진화는 주로 절차적 방식으로 이루어집니다. 유용한 경험은 단순한 저장된 정보가 아니라, 검색, 디버깅 및 검증을 위한 재사용 가능한 절차입니다. 그러나 현재의 평가는 이러한 형태의 전이를 명확하게 측정하지 못합니다. 기존 에이전트 벤치마크는 단일 에피소드의 문제 해결 능력을 평가하는 반면, 메모리 벤치마크는 정보 유지에 초점을 맞추고 절차적 재사용은 고려하지 않습니다. 본 연구에서는 웹 검색, 알고리즘 추론, 소프트웨어 엔지니어링 및 지식 작업의 네 가지 에이전트 영역에서 능력 기반 전송을 통한 에이전트 자기 진화를 평가하기 위한 벤치마크인 EvoAgentBench를 소개합니다. EvoAgentBench는 에이전트 실행 과정에서 추출된 추적 데이터를 기반으로 한 '능력(Abilities)'을 정의하고, 이를 표준화된 운영 단위로 변환하여, 절차적으로 유사한 작업을 연결하는 도메인별 능력 그래프를 구축합니다. 설계상 모든 테스트 작업은 검증된 학습 데이터 측면의 '능력' 지원을 제공합니다. 528/267의 학습/테스트 데이터 분할, 두 가지 구조(scaffold) 및 세 가지 기본 모델(backbone)에 대해, 선별된 '능력' 콘텐츠는 다양한 모델 계열에서 안정적으로 전송되는 것으로 나타났습니다. 그러나 현재까지 자동화된 방법으로 모든 설정에서 긍정적인 성능 향상을 지속적으로 유지하는 데 어려움이 있습니다. EvoAgentBench는 자기 진화 평가를 단순한 정확도 비교에서 벗어나, 경험 인코딩, 라우팅 및 활용에 대한 세밀한 분석으로 전환합니다. 본 벤치마크는 https://huggingface.co/datasets/EverMind-AI/EvoAgentBench 에서 공개적으로 이용할 수 있습니다.

Original Abstract

Agent self-evolution in long-horizon LLM systems is largely procedural: useful experience is not merely stored information, but reusable procedures for searching, debugging, and verification. Yet current evaluations do not isolate this form of transfer. Agent benchmarks test single-episode task solving; memory benchmarks target information retention rather than procedural reuse. We introduce EvoAgentBench, a benchmark for agent self-evolution via Ability-guided transfer across four agentic domains: web research, algorithmic reasoning, software engineering, and knowledge work. EvoAgentBench extracts trace-grounded Abilities from agent executions, canonicalizes them into operational units, and builds domain-specific Ability Graphs linking tasks that share procedural overlap. By design, every test task is backed by verified training-side Ability support. Across a 528/267 train/test split, two scaffolds, and three backbones, curated Ability content transfers reliably across model families, but no current automatic method sustains positive gain in all settings. EvoAgentBench shifts self-evolution evaluation from aggregate accuracy comparison to fine-grained diagnosis of experience encoding, routing, and uptake. The benchmark is publicly available at https://huggingface.co/datasets/EverMind-AI/EvoAgentBench.

3 Citations
1 Influential
27 Altmetric
140.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!