2606.17930v1 Jun 16, 2026 cs.AI

추론 연산량이 최첨단 LLM 평가에 미치는 영향

How Inference Compute Shapes Frontier LLM Evaluation

C. Ududec
C. Ududec
Citations: 589
h-index: 11
H. Coppock
H. Coppock
Citations: 463
h-index: 11
J. McFadyen
J. McFadyen
Citations: 11
h-index: 2
Ole Jorgensen
Ole Jorgensen
Citations: 141
h-index: 3
Kevin Wei
Kevin Wei
Citations: 140
h-index: 3

AI 평가는 도구 사용 및 반복적인 문제 해결을 포함하는 더 복잡한 작업으로 점차 이동하고 있습니다. 그 결과, 성능은 테스트 시점에 사용 가능한 연산량(즉, '추론 연산량')과 이에 대한 할당에 점점 더 민감하게 반응합니다. 그러나 많은 평가에서 여전히 단일하고 제한적인 예산 내에서의 성능만 보고하기 때문에, 낮은 점수는 모델의 실제 능력보다는 평가 설정 자체를 반영할 수 있습니다. 이를 검증하기 위해, 우리는 소프트웨어 엔지니어링, 수학, 의학 및 사이버 보안을 아우르는 7개의 어려운 벤치마크에서 최대 12개의 최첨단 언어 모델을 평가했습니다. 우리는 세 가지 간단한 추론 확장 기법(더 큰 토큰 예산, 컨텍스트 압축 및 반복적인 제출 시도)을 결합하여 통제된 환경에서 실험을 진행했으며, 이는 모델 자체 또는 최소한의 정확성 피드백에 의해 안내됩니다. 주요 결과는 다음과 같습니다. 첫째, 더 큰 토큰 예산은 사이버 보안, FrontierMath, Humanity's Last Exam 및 TerminalBench를 포함한 여러 도메인의 벤치마크에서 성능을 크게 향상시킵니다. 둘째, 고정된 예산을 사용하는 평가는 모델의 발전과 함께 최첨단 모델의 실제 능력을 과소평가할 수 있습니다. 새로운 모델은 더 큰 예산에서 더 높은 성능을 발휘하며, 이를 통해 더욱 어려운 작업을 수행하고 보다 안정적으로 문제를 해결합니다. 셋째, 벤치마크마다 어떤 추론 확장 방법이 가장 효과적인지가 다릅니다. 반복적인 제출 시도는 전반적으로 성능 향상에 도움이 되지만, 더 큰 토큰 예산, 외부 피드백 및 병렬 시도의 가치는 벤치마크에 따라 달라집니다. 전체적으로, 우리의 결과는 벤치마크 점수가 프로토콜에 의존적임을 보여줍니다. 따라서 우리는 평가가 추론 시간 연산량을 기준으로 모델의 능력을 보고하고, 사용된 프로토콜을 명시적으로 밝히며, 특히 안전 또는 정책 관련 환경에서 동일한 예산을 사용하여 광범위한 공유 연산 범위에서 여러 세대의 모델을 비교해야 한다고 주장합니다.

Original Abstract

AI evaluations are shifting toward harder tasks that benefit from longer trajectories involving tool use and iterative problem solving. As a result, performance is increasingly sensitive to the amount and allocation of compute available at test time ("inference compute"). Yet many evaluations still report performance at a single restrictive budget, meaning that low scores may reflect the evaluation setup rather than the model's underlying capability. To test this, we evaluate up to 12 frontier language models on seven challenging benchmarks spanning software engineering, mathematics, medicine, and cybersecurity. We use a controlled setup combining three simple inference-scaling interventions: larger token budgets, context compaction, and repeated submission attempts, guided either by the model itself or by minimal correctness feedback. We find three main results. First, larger token budgets substantially improve performance on benchmarks across multiple domains, including cybersecurity, FrontierMath, Humanity's Last Exam, and TerminalBench. Second, fixed-budget evaluations can increasingly understate frontier capability as models advance. Newer models reach higher performance at large budgets, where they unlock harder tasks and solve them more reliably. Third, benchmarks differ in which inference-scaling methods help most: repeated submission broadly improves performance, but the value of larger token budgets, external feedback, and parallel attempts varies by benchmark. Overall, our results show that benchmark scores are protocol-dependent. We therefore argue that evaluations should report capability as a function of inference-time compute, specify protocol choices explicitly, and compare model generations over a large shared compute range at matched budgets, especially in safety- or policy-relevant settings.

0 Citations
0 Influential
5.5 Altmetric
27.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!