2606.26836v1 Jun 25, 2026 cs.AI

능력 지향 프론티어: 벤치마크는 모델 성능의 82%를 간과합니다.

The Capability Frontier: Benchmarks Miss 82% of Model Performance

Fazl Barez
Fazl Barez
Citations: 1,485
h-index: 17
Narmeen Oozeer
Narmeen Oozeer
Citations: 33
h-index: 3
S. Upadhyay
S. Upadhyay
Citations: 182
h-index: 2
Bradley Fowler
Bradley Fowler
Citations: 2
h-index: 1
Ryan Smith
Ryan Smith
Citations: 0
h-index: 0
Daniel Thi Graviet
Daniel Thi Graviet
Citations: 0
h-index: 0
W. Myers
W. Myers
Citations: 355
h-index: 10
Joshua Greaves
Joshua Greaves
Citations: 2
h-index: 1
A. García
A. García
Citations: 0
h-index: 0
Philip Quirke
Philip Quirke
Citations: 49
h-index: 2
Amirali Abdullah
Amirali Abdullah
Citations: 4
h-index: 1

기존 벤치마크는 일반적으로 단일 모델을 단 한 번 실행하여 정확도를 보고합니다. 이는 실제 LLM(대규모 언어 모델)의 능력을 체계적으로 과소평가하며, 특히 데이터 분포가 이질적인 경우 더욱 그렇습니다. (i) 각 모델은 특성에 따라 다른 질문에 대해 올바른 답변을 제공하고, (ii) 주어진 예산 내에서 여러 번의 생성을 통해 최적의 결과를 선택할 수 있습니다. 이러한 격차를 정량화하기 위해, 우리는 '능력 지향 프론티어(Capability Frontier)'라는 개념을 도입합니다. 이는 모델 집합에 대한 파레토 프론티어로, 각 비용 수준에서 최적으로 모델과 생성을 선택했을 때 달성 가능한 최고의 성능을 나타냅니다. 우리의 방법은 두 가지 상반된 편향을 보정합니다: 단일 모델 평가로 인한 과소평가와, 잡음이 많은 샘플의 최대값을 취함으로써 발생하는 과대평가입니다. 우리는 코딩, 추론, 의학, 사실성, 지시 따르기 및 에이전트 작업 등 16개의 광범위하게 사용되는 벤치마크에서 21개의 LLM을 연구하고, 각 벤치마크의 최고 성능 모델과 동일한 비용 수준에서의 능력 지향 프론티어 성능을 비교합니다. 단일 모델 평가를 보정한 결과, 오류율이 54% 감소했습니다. 또한 단일 실행을 보정한 결과, 전체적으로 82%의 개선 효과가 나타났으며, SOTA(State-of-the-Art, 최첨단) 정확도를 85% 더 낮은 비용으로 달성할 수 있었습니다. 이러한 실증적 결과를 뒷받침하기 위해, 우리는 제어된 확률 시뮬레이션을 통해 검색 질문 주제의 엔트로피가 높을수록 오라클 라우팅과 최고의 단일 모델 간의 성능 격차가 거의 단조적으로 증가한다는 것을 보여줍니다. 우리의 연구 결과는 LLM의 집단적인 능력이 크게 과소평가되어 왔으며, 이는 데이터 이질성 및 다중 도메인 환경에서의 평가 및 배포에 중요한 함의를 갖는다는 것을 시사합니다.

Original Abstract

Existing benchmarks typically report accuracy for a single model on a single run. This systematically understates real-world LLM capabilities, particularly under heterogeneous data distributions: (i) different models get different questions correct according to their specializations, and (ii) given a budget, multiple generations can be sampled and selectively retained. To quantify this gap, we introduce the Capability Frontier: a Pareto frontier over a set of models that characterizes the best achievable performance at each cost level under optimal selection across models and generations (i.e., via an oracle). Our construction corrects for two opposing biases: underestimation from single-model evaluation and overestimation from taking maxima over noisy samples. We study 21 LLMs across 16 widely used benchmarks spanning coding, reasoning, medicine, factuality, instruction following, and agentic tasks, comparing Capability Frontier performance at matched cost to each benchmark's top-performing model. Correcting for single-model evaluation yields a 54% error rate reduction; additionally correcting for single runs yields an 82% improvement, with SOTA accuracy matched at 85% cost reduction. Complementing these empirical results, we use controlled probabilistic simulations to show that higher query topic entropy produces a near-monotonic increase in the performance gap between oracle routing and the best single model. Our findings suggest collective LLM capabilities are substantially underestimated, with implications for evaluation and deployment in data-heterogeneous, multi-domain settings.

1 Citations
0 Influential
8.5 Altmetric
43.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!