2607.12338v1 Jul 14, 2026 cs.AI

에이전트 성능 평가를 위한 충분한 작업량은 얼마나 될까? 공개 LLM 에이전트 벤치마크 재분석

How Many Tasks Are Enough for Agent Benchmark Decisions? A Replay Analysis of Public LLM Agent Benchmarks

Wei-Jung Huang
Wei-Jung Huang
Citations: 0
h-index: 0

에이전트 벤치마크는 일반적으로 모든 작업이 완료된 후에 두 에이전트를 비교하지만, 비용 때문에 부분적인 실행을 고려하게 되는 경우가 있습니다. 작업 비율만으로는 부분 실행이 전체 벤치마크와 동일한 결론을 도출하는지 알 수 없습니다. 본 연구에서는 SWE-bench, AppWorld, tau-bench의 공개된 작업 레코드를 재사용하여 이 문제를 분석합니다. 제한된 예산 내에서 부분 실행이 유효하려면, 전체 벤치마크의 결정과 일치해야 하며, 필요한 작업 그룹을 포함해야 하고, 목표 비율 이상의 비교가 해결되지 않아야 합니다. 필요한 작업 비율은 현저하게 다릅니다. 엄격한 기준(5% 범위 내에서 0% 증가) 하에서, AppWorld는 15%, tau-bench는 25%, SWE-bench Verified는 90%에서 모든 목표를 충족합니다. 반면, SWE-bench Lite는 주요 규칙을 적용했을 때 95%까지도 모든 목표를 충족하지 못합니다. 부분 평가 보고서에는 한 에이전트가 다른 에이전트보다 얼마나 더 뛰어나야 하는지, 어떤 작업이 선택되었는지, 어떤 커버리지 규칙이 필요한지, 어떤 결정 규칙이 사용되는지, 그리고 몇 개의 비교가 해결되지 않은 상태로 남을 수 있는지 명시해야 합니다.

Original Abstract

Agent benchmarks often compare two agents after all tasks have run, but costly evaluations make partial runs tempting. A task fraction alone does not show whether a partial run supports the same pairwise conclusion as the completed benchmark. We study this question by replaying completed public task-level records from SWE-bench, AppWorld, and tau-bench. A partial budget counts as enough only when it supports the completed benchmark's decision, covers required task groups, and leaves no more than a target fraction of comparisons unresolved. The required task fraction varies sharply. At the strict 0 percentage point threshold on a 5 percentage point budget grid, AppWorld first meets all targets at 15 percent, tau-bench at 25 percent, and SWE-bench Verified at 90 percent; SWE-bench Lite does not meet all targets by 95 percent under the primary coverage rule. Partial-evaluation reports should state how much one agent must outperform another, how tasks are selected, what coverage rule is required, what decision rule is used, and how many comparisons may remain unresolved.

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!