2608.06144v1 Aug 06, 2026 cs.AI

FinEvo-Bench: 전문가 금융 워크플로우에서 스스로 진화하는 에이전트를 위한 장기적 벤치마크

FinEvo-Bench: A Longitudinal Benchmark for Self-Evolving Agents in Professional Financial Workflows

Renzhao Liang
Renzhao Liang
Citations: 6
h-index: 1
Chenggang Xie
Chenggang Xie
Citations: 8
h-index: 1
Lifan Guo
Lifan Guo
Citations: 69
h-index: 4
Feng Chen
Feng Chen
Citations: 39
h-index: 3
Kang Zhou
Kang Zhou
Citations: 75
h-index: 3
Bo Deng
Bo Deng
Citations: 22
h-index: 1
Chongyang Tao
Chongyang Tao
Citations: 3
h-index: 1
Chi Zhang Beihang University
Chi Zhang Beihang University
Citations: 0
h-index: 0
Qwen Team
Qwen Team
Citations: 94
h-index: 2
Alibaba Cloud Computing
Alibaba Cloud Computing
Citations: 0
h-index: 0
Xuanren Chen
Xuanren Chen
Citations: 0
h-index: 0

대부분의 에이전트 벤치마크는 작업을 독립적으로 평가하며, 하나의 작업에서의 경험이 이후 작업에 얼마나 도움이 되는지 측정하지 못합니다. 기존의 자기 진화 벤치마크는 전문적인 워크플로우, 개방형 결과물 및 다각적 평가를 종합적으로 다루지 않습니다. 본 연구에서는 120개의 실제 사례 기반 작업과 여섯 가지 금융 분야에 걸쳐 20개의 비즈니스 시나리오를 포함하는 장기 벤치마크인 FinEvo-Bench를 소개합니다. 기관에서 제공하는 전문적인 절차가 필요한 작업과 제약을 정의하며, 기관에서 제공하고 공개적으로 문서화된 사례들이 작업 사실을 구성합니다. 각 시나리오는 관련된 여섯 가지 사례로 구성되며, 이들은 동일한 전문적인 절차와 수동으로 검토된 작업 품질 및 금융 규정 준수 기준을 공유합니다. 동일한 Qwen3.7-Max 백본을 사용하는 네 가지 자기 진화 에이전트 모델을 비교했으며, 세 개의 독립적으로 섞인 글로벌 작업 스트림을 사용했습니다. 쌍을 이루는 비진화 제어 그룹은 각 모델의 유지된 경험으로부터 얻는 자기 진화 효과를 추정하며, Claude Opus 4.6 기반의 독립적인 Claude Code 평가 에이전트는 모든 결과를 평가합니다. Letta 모델은 가장 높은 진화 점수(91.65)와 가장 적은 규정 위반 건수(작업당 0.09건)를 달성했으며, Codex 모델은 가장 큰 자기 진화 향상(+19.37)을 보였습니다. 모든 모델에서 진화 조건은 점수를 9.33~19.37점 높이고 작업당 규정 위반 건수를 0.12~0.44건 줄였습니다. 시나리오 내의 순위 4~6에서의 쌍을 이루는 점수 향상은 순위 1~3에서의 향상보다 6.10~8.70점 더 높았습니다. Claude Code에서 기술만으로 진화하는 모델은 기억 기반 또는 혼합 기억-기술 진화 모델보다 더 높은 작업 품질과 규정 준수 건수를 보였습니다. 모든 네 가지 모델에서 루브릭 피드백은 참조 답변 피드백보다 더 높은 점수와 규정 위반 건수를 제공했습니다. FinEvo-Bench는 전문적인 성능과 자기 진화 능력 모두를 측정합니다. 즉, 에이전트가 이전 경험을 활용하여 후속 개선을 얼마나 효과적으로 수행하는지를 평가합니다.

Original Abstract

Most agent benchmarks evaluate tasks independently and cannot measure whether experience from one task helps with later tasks. Existing self-evolution benchmarks do not jointly cover professional workflows, open-ended deliverables, and multi-aspect evaluation. We introduce FinEvo-Bench, a longitudinal benchmark with 120 real-case-grounded tasks, 20 business scenes across six financial domains. Institution-provided professional procedures define the required operations and constraints. Eligible institution-provided and publicly documented cases supply the task facts. Each scene contains six related but substantively distinct cases that share a professional procedure and a manually reviewed rubric for task quality and financial compliance. We compare four self-evolving agent scaffolds using the same Qwen3.7-Max backbone and three independently shuffled, globally interleaved task streams. Paired non-evolving controls estimate each scaffold's self-evolution gain from retained experience, while an independent Claude Code scoring agent backed by Claude Opus 4.6 evaluates all outputs. Letta achieves the highest evolved score (91.65) and fewest compliance issues (0.09 per task); Codex achieves the largest self-evolution gain (+19.37). Across scaffolds, the evolving condition raises scores by 9.33-19.37 points and reduces compliance issues by 0.12-0.44 per task. Paired score gains at within-scene ranks 4-6 exceed those at ranks 1-3 by 6.10-8.70 points. In Claude Code, skill-only evolution produces higher task quality and fewer compliance issues than memory-only and combined memory-skill evolution. Across all four scaffolds, rubric feedback also yields higher scores and fewer compliance issues than reference-answer feedback. FinEvo-Bench measures both professional performance and self-evolution ability: how effectively an agent turns prior experience into later improvement.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!