2605.00300v1 May 01, 2026 cs.AI

토큰 아레나: AI 추론에서 에너지와 인지 능력을 통합하는 지속적인 벤치마크

Token Arena: A Continuous Benchmark Unifying Energy and Cognition in AI Inference

Yuxuan Gao
Yuxuan Gao
Citations: 6
h-index: 1
Yixu Yu
Yixu Yu
Citations: 0
h-index: 0
Megan Wang
Megan Wang
Citations: 6
h-index: 1

기존의 공개 추론 벤치마크는 AI 시스템을 모델 및 서비스 제공업체 수준에서 비교하지만, 실제 배포 결정은 엔드포인트 수준에서 이루어집니다. 엔드포인트는 특정 양자화, 디코딩 전략, 지역, 그리고 서비스 스택이 적용된 (서비스 제공업체, 모델, SKU) 튜플을 의미합니다. 본 논문에서는 추론을 엔드포인트 수준의 세분화로 측정하고, 다섯 가지 핵심 요소(출력 속도, 첫 번째 토큰까지의 시간, 워크로드 기반 가격, 효과적인 컨텍스트, 그리고 실제 엔드포인트에서의 품질)를 통합하여 세 가지 주요 지표(정답 1개당 소비 전력, 정답 1개당 비용, 그리고 엔드포인트의 충실도(출력 분포와 자체 참조 값과의 유사성))를 제시하는 지속적인 벤치마크인 '토큰 아레나'를 소개합니다. 본 프레임워크의 혁신성은 실증적이고 방법론적입니다. 78개의 엔드포인트와 12개의 모델 패밀리를 대상으로 실험한 결과, 동일한 모델이라도 엔드포인트에 따라 수학 및 코딩 문제에 대한 평균 정확도가 최대 12.5점 차이, 자체 참조 값과의 유사성이 최대 12점 차이, 최대 지연 시간이 최대 10배 차이, 그리고 모델링된 정답 1개당 소비 전력이 최대 6.2배 차이를 보입니다. 또한, 워크로드에 따른 가격 조합은 순위표를 크게 변화시킵니다. 챗(3:1 입력:출력) 설정에서 상위 10개 엔드포인트 중 7개가 검색 증강(20:1) 설정에서는 순위 밖으로 밀려났으며, 추론 설정(1:5)은 가격 면에서 불리한 점이 있는 최첨단 폐쇄 모델을 순위 상승시킵니다. 본 프레임워크, 스키마, 측정 및 평가 도구, 그리고 v1.0 순위표 스냅샷을 CC BY 4.0 라이선스로 공개합니다. '토큰 아레나'는 단일 순위가 아닌 방법론이며, 모든 출처 및 제한 사항을 공개하고 외부 재현을 환영합니다.

Original Abstract

Public inference benchmarks compare AI systems at the model and provider level, but the unit at which deployment decisions are actually made is the endpoint: the (provider, model, stock-keeping-unit) tuple at which a specific quantization, decoding strategy, region, and serving stack is exposed. We introduce TokenArena, a continuous benchmark that measures inference at endpoint granularity along five core axes (output speed, time to first token, workload-blended price, effective context, and quality on the live endpoint) and synthesizes them, together with a modeled energy estimate, into three headline composites: joules per correct answer, dollars per correct answer, and endpoint fidelity (output-distribution similarity to a first-party reference). The framework's novelty is empirical and methodological. Across 78 endpoints serving 12 model families, the same model on different endpoints differs in mean accuracy by up to 12.5 points on math and code, in fingerprint similarity to first party by up to 12 points, in tail latency by an order of magnitude, and in modeled joules per correct answer by a factor of 6.2. We further show that workload-aware blended pricing reorders the leaderboard substantially: 7 of 10 top-ranked endpoints under the chat preset (3:1 input:output) fall out of the top 10 under the retrieval-augmented preset (20:1), and the reasoning preset (1:5) elevates frontier closed models that the chat preset penalizes on price. We release the framework, schema, probe and eval harness, and a v1.0 leaderboard snapshot under CC BY 4.0. TokenArena is a methodology, not a single ranking; we publish full provenance and limitations and welcome external replication.

0 Citations
0 Influential
0.5 Altmetric
2.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!