2602.16763v1 Feb 18, 2026 cs.AI

AI 벤치마크가 정체될 때: 벤치마크 포화 현상에 대한 체계적인 연구

When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

Leshem Choshen
Leshem Choshen
Citations: 250
h-index: 9
Mrinmaya Sachan
Mrinmaya Sachan
Citations: 5,987
h-index: 43
M. Kochenderfer
M. Kochenderfer
Citations: 2,206
h-index: 24
Siddhesh Pawar
Siddhesh Pawar
Citations: 177
h-index: 4
Mubashara Akhtar
Mubashara Akhtar
Citations: 346
h-index: 9
Anka Reuel
Anka Reuel
Citations: 1,295
h-index: 14
Prajna Soni
Prajna Soni
Citations: 26
h-index: 3
Sanchit Ahuja
Sanchit Ahuja
Citations: 254
h-index: 8
Pawan Sasanka Ammanamanchi
Pawan Sasanka Ammanamanchi
Citations: 3,660
h-index: 8
Ruchit Rawal
Ruchit Rawal
Citations: 206
h-index: 6
Vilém Zouhar
Vilém Zouhar
Citations: 109
h-index: 7
Chenxi Whitehouse
Chenxi Whitehouse
Citations: 800
h-index: 13
Dayeon Ki
Dayeon Ki
Citations: 575
h-index: 6
Jennifer Mickel
Jennifer Mickel
Citations: 218
h-index: 5
Marek vSuppa
Marek vSuppa
Citations: 185
h-index: 3
Jan Batzner
Jan Batzner
Citations: 156
h-index: 7
Jenny Chim
Jenny Chim
Citations: 13
h-index: 2
Jeba Sania
Jeba Sania
Citations: 15
h-index: 2
Yanan Long
Yanan Long
Citations: 104
h-index: 4
Hossein A. Rahmani
Hossein A. Rahmani
Citations: 13
h-index: 2
Christina Q. Knight
Christina Q. Knight
Citations: 57
h-index: 5
Yiyang Nan
Yiyang Nan
Citations: 220
h-index: 7
Yu Fan
Yu Fan
Citations: 61
h-index: 3
Shubham Singh
Shubham Singh
Citations: 36
h-index: 4
Subramanyam Sahoo
Subramanyam Sahoo
Citations: 20
h-index: 2
Eliya Habba
Eliya Habba
Citations: 61
h-index: 5
Usman Gohar
Usman Gohar
Iowa State University
Citations: 445
h-index: 8
Robert Scholz
Robert Scholz
Citations: 30
h-index: 3
Arjun Subramonian
Arjun Subramonian
Meta FAIR
Citations: 4,170
h-index: 15
Jingwei Ni
Jingwei Ni
ETH Zürich
Citations: 541
h-index: 12
Sanmi Koyejo
Sanmi Koyejo
Citations: 4,387
h-index: 25
Stella Biderman
Stella Biderman
Citations: 217
h-index: 4
Z. Talat
Z. Talat
Citations: 8
h-index: 1
Irene Solaiman
Irene Solaiman
Citations: 4,597
h-index: 9
Srishti Yadav
Srishti Yadav
Citations: 173
h-index: 5
Avijit Ghosh
Avijit Ghosh
Citations: 143
h-index: 6
Jyoutir Raj
Jyoutir Raj
Citations: 11
h-index: 2

인공지능(AI) 벤치마크는 모델 개발의 진전을 측정하고 배포 결정을 안내하는 데 중요한 역할을 합니다. 그러나 많은 벤치마크는 빠르게 포화 상태에 도달하여, 더 이상 최고 성능 모델을 구분할 수 없게 되어 장기적인 가치를 감소시킵니다. 본 연구에서는 주요 모델 개발사의 기술 보고서에서 선정한 60개의 대규모 언어 모델(LLM) 벤치마크에 대한 벤치마크 포화 현상을 분석합니다. 벤치마크 포화를 유발하는 요인을 파악하기 위해, 작업 설계, 데이터 구성 및 평가 형식을 포괄하는 14가지 속성을 기준으로 벤치마크를 특성화했습니다. 각 속성이 포화율에 어떻게 기여하는지 살펴보는 5가지 가설을 검증했습니다. 분석 결과, 거의 절반의 벤치마크가 포화 상태를 보이는 것으로 나타났으며, 벤치마크의 노후화될수록 포화율이 증가했습니다. 주목할 점은 테스트 데이터의 공개 여부(예: 공개 vs. 비공개)가 포화 현상을 막는 효과가 없으며, 전문가가 선별한 벤치마크는 크라우드소싱된 벤치마크보다 포화 현상에 더 잘 저항한다는 것입니다. 본 연구 결과는 벤치마크의 수명을 연장하는 설계 선택을 강조하고, 보다 지속 가능한 평가 전략을 수립하는 데 필요한 정보를 제공합니다.

Original Abstract

Artificial intelligence benchmarks are an important mechanism to measure model progress and guide deployment decisions. However, benchmarks quickly "saturate", making it difficult to differentiate models and diminishing their long-term value. In this study, we define benchmark saturation and analyze it across 60 language model benchmarks using 14 properties that relate to saturation. We find that nearly half of our benchmarks exhibit saturation, with rates increasing with age. Further, we find that resilience to saturation is impacted by expert-curation, not by public test data. Our results suggest that design choices can extend benchmark longevity and inform more durable evaluation approaches.

18 Citations
0 Influential
21.5 Altmetric
125.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!