2606.19704v1 Jun 18, 2026 cs.AI

정적인 순위표를 넘어서: LLM 에이전트 평가를 위한 예측 타당성

Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents

Yihang Sun
Yihang Sun
Citations: 63
h-index: 3
K. E. Maghraoui
K. E. Maghraoui
Citations: 1,479
h-index: 18
Rui Li
Rui Li
Citations: 7
h-index: 1
Sagar Kumar
Sagar Kumar
Citations: 104
h-index: 1
Dhaval Patel
Dhaval Patel
Citations: 31
h-index: 3
Tianyang Xu
Tianyang Xu
Citations: 10
h-index: 1
Tianjun Feng
Tianjun Feng
Citations: 45
h-index: 1
C. Tsai
C. Tsai
Citations: 10
h-index: 1
Shuxin Lin
Shuxin Lin
Citations: 92
h-index: 5
Yusheng Li
Yusheng Li
Citations: 2
h-index: 1
Wei Xin
Wei Xin
Citations: 7
h-index: 1
A. Bhandari
A. Bhandari
Citations: 11
h-index: 2
Tanisha Rathod
Tanisha Rathod
Citations: 2
h-index: 1
Aaron Fan
Aaron Fan
Citations: 32
h-index: 3
S. Shejwal
S. Shejwal
Citations: 8
h-index: 1
Tomas Pasiecznik
Tomas Pasiecznik
Citations: 0
h-index: 0
Tanmay Agarwal
Tanmay Agarwal
Citations: 1,165
h-index: 6
Rohith Kanathur
Rohith Kanathur
Citations: 2
h-index: 1
Sam Colman
Sam Colman
Citations: 0
h-index: 0
A. Sheikh
A. Sheikh
Citations: 17
h-index: 2
D. Bahl
D. Bahl
Citations: 1
h-index: 1
Anny Li
Anny Li
Citations: 0
h-index: 0
Krish Veera
Krish Veera
Citations: 9
h-index: 2
A. Merchant
A. Merchant
Citations: 0
h-index: 0
S. Bhure
S. Bhure
Citations: 0
h-index: 0
Sajal Kumar Goyla
Sajal Kumar Goyla
Citations: 0
h-index: 0
Chengrui Li
Chengrui Li
Citations: 0
h-index: 0
Kirthana Natarajan
Kirthana Natarajan
Citations: 0
h-index: 0
T. Ajai
T. Ajai
Citations: 0
h-index: 0
Rujing Li
Rujing Li
Citations: 0
h-index: 0
Vivek Iyer
Vivek Iyer
Citations: 2
h-index: 1
S. Vijayakumar
S. Vijayakumar
Citations: 2
h-index: 1
Yitong Bai
Yitong Bai
Citations: 3
h-index: 1
Ayal Yakobe
Ayal Yakobe
Citations: 1
h-index: 1
Darief Maes
Darief Maes
Citations: 0
h-index: 0
Yassine Jebbouri
Yassine Jebbouri
Citations: 0
h-index: 0
Thai On
Thai On
Citations: 32
h-index: 2
Vera M Mazeeva
Vera M Mazeeva
Citations: 0
h-index: 0
Winston Li
Winston Li
Citations: 0
h-index: 0
Yuval Shemla
Yuval Shemla
Citations: 6
h-index: 1
Yeshitha Bhuvanesh
Yeshitha Bhuvanesh
Citations: 0
h-index: 0
Rushin Bhatt
Rushin Bhatt
Citations: 5
h-index: 1
Siddharth Chethan Gowda
Siddharth Chethan Gowda
Citations: 0
h-index: 0
Alisha Vinod
Alisha Vinod
Citations: 0
h-index: 0
C. Cahill
C. Cahill
Citations: 4
h-index: 1
Shriya Aishani Rachakonda
Shriya Aishani Rachakonda
Citations: 0
h-index: 0
Yun Chen
Yun Chen
Citations: 80
h-index: 3
A. Agrawal
A. Agrawal
Citations: 0
h-index: 0
Aman Upganlawar
Aman Upganlawar
Citations: 3
h-index: 1
Mao Le Jonathan Ang
Mao Le Jonathan Ang
Citations: 0
h-index: 0
Yubin Go
Yubin Go
Citations: 9
h-index: 2
Madhav Rajkondawar
Madhav Rajkondawar
Citations: 0
h-index: 0
Yang Chen
Yang Chen
Citations: 58
h-index: 4
Trish Maturi
Trish Maturi
Citations: 12
h-index: 2
Ananya Kapoor
Ananya Kapoor
Citations: 6
h-index: 2
Andrew Li
Andrew Li
Citations: 0
h-index: 0
Shrey Arora
Shrey Arora
Citations: 3
h-index: 1
Mana Abbaszadeh
Mana Abbaszadeh
Citations: 0
h-index: 0
Shenlin Li
Shenlin Li
Citations: 0
h-index: 0
Charles Xu
Charles Xu
Citations: 0
h-index: 0
Byeolah Kwon
Byeolah Kwon
Citations: 0
h-index: 0

에이전트 벤치마크는 빠르게 증가하고 있지만, 단일 벤치마크가 실제 배포 환경에서 드러나는 다양한 측면 중 네다섯 개 이상을 포괄하는 경우는 드뭅니다. 본 논문에서는 현재까지 가장 광범위하고 심층적인 MCP 기반 산업용 에이전트 벤치마크 연구 결과를 종합적으로 분석합니다. 이 연구는 새로운 자산 클래스(멀티모달 시각 정보 포함), 다양한 오케스트레이션 방식, 검색 전략, 추론 모드, 인프라 최적화 및 평가 방법론에 대한 14개의 병렬 구현 연구를 다룹니다. 이러한 연구와 함께 기존의 7가지 에이전트 벤치마크 결과를 종합하여 분석한 결과, 집계 점수를 기반으로 한 순위표는 실제 배포 환경에서 에이전트를 평가하는 데 있어 중요한 정보를 누락시킬 수 있음을 주장합니다. 집계 점수에 따른 순위는 일반화되지 않은 환경에서도 그대로 유지되지 않으며, 최근 공개된 경쟁 대회 결과를 분석한 결과 이러한 순위 불안정성에 대한 직접적인 경험적 증거를 제시합니다. 본 연구에서는 예측 타당성, 즉 학습 데이터에서의 순위와 실제 데이터에서의 순위 간의 상관관계를 기준으로 에이전트 평가 시스템을 구성하는 방법을 제안하고, HELM 및 후속 모델에서 다루는 배포 관련 주요 측면들을 측정할 수 있는 12단계 측정 도구를 제시합니다. 이 방법론은 세 가지 검증 가능한 일반화 기준과 명확한 임계값을 통해 구현되며, 현재까지의 증거는 이를 부분적으로 뒷받침하지만, 아직 충분히 확정하기에는 부족합니다. 마지막으로, 본 논문에서는 사전 등록된 파일럿 연구 설계와 향후 에이전트 벤치마크가 제공해야 할 내용에 대한 비전을 제시합니다.

Original Abstract

Agent benchmarks are growing fast, but no single benchmark touches more than four or five of the dimensions that deployment exposes. This paper aggregates the largest coordinated deep-dive of one MCP-based industrial-agent benchmark to date: fourteen parallel implementation studies covering new asset classes (including a multi-modal visual extension), alternative orchestrations, retrieval strategies, reasoning modes, infrastructure optimizations, and evaluation-methodology probes. Consolidating those studies with seven prior agent benchmarks, we argue that aggregate-score leaderboards systematically underspecify deployed-agent evaluation. Rankings derived from aggregate scores do not transfer to out-of-distribution settings; recent public-to-hidden competition retrospectives provide direct empirical evidence of this rank instability. We propose ranking configurations by predictive validity, the correlation between in-sample and out-of-sample rank, rather than in-sample mean, and report a twelve-tier measurement apparatus that exposes the deployment-relevant dimensions HELM and its agent-era successors collapse. The position is operationalized through three falsifiable out-of-distribution criteria with explicit thresholds; existing evidence partly supports it but is too thin to confirm. We close with a pre-registered pilot design and a field-level vision for what the next generation of agentic benchmarks should report.

0 Citations
0 Influential
9 Altmetric
45.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!