2606.05670v1 Jun 04, 2026 cs.AI

더 많은 에이전트가 도움이 되는가? LLM 에이전트 워크플로우의 통제되고 프로토콜 기반 평가

Do More Agents Help? Controlled and Protocol-Aligned Evaluation of LLM Agent Workflows

Jiaqi Shao
Jiaqi Shao
Citations: 89
h-index: 4
Bing Luo
Bing Luo
Citations: 83
h-index: 4
Yuhan Fu
Yuhan Fu
Citations: 2
h-index: 1
Ruishan Fang
Ruishan Fang
Citations: 6
h-index: 2
Tao Lin
Tao Lin
Citations: 28
h-index: 2
Huiyuan Zheng
Huiyuan Zheng
Citations: 109
h-index: 3
Zhengtao Zhu
Zhengtao Zhu
Citations: 0
h-index: 0

동일한 벤치마크 로더, 도구 접근 권한, 답변 형식, 사용량 기록 및 실행 경로 로깅을 공유하는 시스템에서, 더 많은 에이전트를 추가하면 LLM 워크플로우에 도움이 되는가? 본 연구에서는 BenchAgent라는 평가 프레임워크를 소개하며, 이 프레임워크는 단일 에이전트, 고정 다중 에이전트(MAS) 및 진화하는 MAS 워크플로우를 하나의 표준화된 실행 및 로깅 프로토콜 하에 배치합니다. BenchAgent는 GPT-4.1을 사용하여 10개의 추론, 코딩 및 도구 사용 벤치마크에서 이러한 내부 워크플로우를 평가하고, 별도로 런타임 생성 워크플로우에 대한 프로토콜 기반 외부(PAE) GAIA 연구 결과를 보고합니다. SI 조건 하에서 테스트된 6개의 MAS 중 최대 하나만이 벤치마크 균형 평균 정확도 측면에서 해당 단일 에이전트 기준을 능가하며, EvoAgent는 Wilson의 단일 실행 지침 범위 내에 있습니다. 나머지 5개는 2.56~11.29점 뒤처지며, 더 높은 정확도를 얻기 위한 비용 대비 효율성이 낮습니다. PAE GAIA 스냅샷에서 Claude-Code 스타일의 런타임 워크플로우는 전체적으로 66.72%의 성능을 보이며, Level 3에서는 69.23%의 성능을 보여, 가장 강력한 비-Claude 기반 모델인 Jarvis (고정 MAS)보다 20점 이상 높은 성능을 나타냅니다.

Original Abstract

Does adding more agents help an LLM workflow once compared systems share the same benchmark loader, tool access, answer contract, usage accounting, and trajectory logging? We introduce BenchAgent, an evaluation framework that places single-agent, fixed multi-agent (MAS), and evolving MAS workflows under one normalized execution and logging protocol. BenchAgent evaluates these substrate-internal workflows across ten reasoning, coding, and tool-use benchmarks with GPT-4.1, and separately reports a Protocol-Aligned External (PAE) GAIA study of a runtime-generated workflow. Under SI conditions, at most one of six tested MAS exceeds the matched single-agent anchor on benchmark-balanced average accuracy: EvoAgent lies within the Wilson one-run guidance, while the remaining five trail by 2.56-11.29 points and occupy more expensive accuracy-cost trade-offs. On the PAE GAIA snapshot, a Claude-Code-style runtime workflow reaches 66.72% overall and 69.23% on Level 3, more than 20 points above the strongest non-Claude baseline, Jarvis, a fixed MAS.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!