2607.01211v1 Jul 01, 2026 cs.SE

성능 최적화 벤치마크가 코딩 에이전트의 성능을 신뢰성 있게 측정하는가?

Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents?

Yuling Shi
Yuling Shi
Citations: 398
h-index: 11
Zhi Chen
Zhi Chen
Citations: 52
h-index: 4
Zhensu Sun
Zhensu Sun
Citations: 536
h-index: 11
David Lo
David Lo
Citations: 13
h-index: 2
Lingxiao Jiang
Lingxiao Jiang
Citations: 292
h-index: 6

GSO, SWE-Perf, SWE-fficiency와 같은 저장소 수준의 성능 최적화 벤치마크는 실제 저장소에 패치를 적용하고 실행 시간을 최적화되지 않은 기준과 공식 참조 패치와 비교하여 코딩 에이전트를 평가합니다. 이러한 벤치마크의 순위 점수는 코딩 에이전트 발전의 지표로 점점 더 많이 사용되고 있지만, 해당 점수는 실행 시간의 불안정성, 벤치마크별 채점 규칙, 그리고 이미 최소한 하나의 공개 제출물이 해결한 작업 수를 혼동할 수 있습니다. 본 연구에서는 세 가지 벤치마크에 대한 이러한 문제점을 분석합니다. 첫째, Google Cloud에서 사용되는 네 가지 일반적인 머신 유형에서 740개의 코드 최적화 작업을 대상으로 공식 참조 패치를 재실행했습니다. 대부분의 벤치마크 작업은 재실행 가능하지만, 참조 패치가 모든 머신에서 원래 벤치마크 유효성 규칙을 만족하는 경우는 GSO의 경우 102개 작업 중 39개, SWE-Perf의 경우 140개 작업 중 11개, SWE-fficiency의 경우 498개 작업 중 411개에 불과했습니다. 특히 SWE-Perf는 많은 참조 패치가 실행 시간 변화를 거의 발생시키지 않는다는 점에서 불안정성이 높습니다. 둘째, 공개 제출물의 순위가 벤치마크 채점 규칙에 크게 의존한다는 것을 보여줍니다. GSO와 SWE-fficiency에서 공유된 여덟 개의 공개 제출물 중, 공식 순위는 28개의 쌍별 제출물 비교 중 9건에서 의견이 일치하지 않았으며, SWE-fficiency의 리더보드 채점 규칙은 최악의 10개 작업에 대해 과도하게 높은 점수 가중치(58.5%~82.8%)를 부여합니다. 셋째, 각 작업별로 10개의 공개 제출물을 살펴보면, 재실행이 가능한 GSO 및 SWE-fficiency 작업의 85.3%(384/450)에서 최소 하나의 제출물이 참조 패치와 일치하거나 능가하며, 최적화되지 않은 기준 코드보다 99.8%(449/450)에서 더 나은 성능을 보입니다. 본 연구는 리더보드 점수 외에도 더욱 신뢰할 수 있는 성능 지표를 가진 작업들을 식별하고, 작업별 점수 기여도를 정량화하며, 집계 순위로 인해 숨겨진 남은 성능 격차를 드러냄으로써 기존 연구를 보완합니다.

Original Abstract

Repository-level performance-optimization benchmarks such as GSO, SWE-Perf and SWE-fficiency evaluate coding agents by applying patches to real repositories and comparing runtime against unoptimized baselines and official reference patches. Their leaderboard scores are increasingly used as evidence of coding-agent progress, but those scores can conflate runtime instability, benchmark-specific scoring rules, and how many tasks are already solved by at least one public submission. We audit these issues across the three benchmarks. First, we replay the official reference patches for 740 code optimization tasks across four common types of Google Cloud machines. Most benchmark tasks can be replayed, but their reference patches satisfy the original benchmark validity rules in every cross-machine replay for only 39/102 GSO tasks, 11/140 SWE-Perf tasks, and 411/498 SWE-fficiency tasks; SWE-Perf is especially fragile because many reference patches produce close-to-zero runtime changes. Second, we show that public submission rankings depend strongly on the benchmark scoring rule. Among eight public submissions shared by GSO and SWE-fficiency, the official rankings disagree on 9 of 28 pairwise submission comparisons, and SWE-fficiency's leaderboard scoring rule assigns the worst ten tasks overly high score weights of 58.5%-82.8%. Third, looking across 10 public submissions for each task, we find that at least one submission matches or beats the reference patch on 85.3% (384/450) of replay-valid GSO and SWE-fficiency tasks, and beats the unoptimized base code on 99.8% (449/450). Our study complements leaderboard scores by identifying tasks with more reliable performance signals, quantifying per-task score contributions, and exposing the remaining performance gaps that are hidden by aggregate rankings.

0 Citations
0 Influential
5.5 Altmetric
27.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!