2606.09809v1 Jun 08, 2026 cs.AI

평가 카드: AI 평가 보고를 위한 해석 레이어

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting

Leshem Choshen
Leshem Choshen
Citations: 250
h-index: 9
Avijit Ghosh
Avijit Ghosh
Citations: 44
h-index: 4
M. Kochenderfer
M. Kochenderfer
Citations: 2,206
h-index: 24
David Manheim
David Manheim
Citations: 59
h-index: 3
Anka Reuel
Anka Reuel
Citations: 1,295
h-index: 14
Jennifer Mickel
Jennifer Mickel
Citations: 218
h-index: 5
Jan Batzner
Jan Batzner
Citations: 156
h-index: 7
Jenny Chim
Jenny Chim
Citations: 13
h-index: 2
Jeba Sania
Jeba Sania
Citations: 15
h-index: 2
Yanan Long
Yanan Long
Citations: 104
h-index: 4
Eliya Habba
Eliya Habba
Citations: 61
h-index: 5
Usman Gohar
Usman Gohar
Iowa State University
Citations: 445
h-index: 8
Sanmi Koyejo
Sanmi Koyejo
Citations: 4,387
h-index: 25
Stella Biderman
Stella Biderman
Citations: 217
h-index: 4
Irene Solaiman
Irene Solaiman
Citations: 4,597
h-index: 9
Asaf Yehudai
Asaf Yehudai
Citations: 407
h-index: 9
Srishti Yadav
Srishti Yadav
Citations: 173
h-index: 5
Michael Hardy
Michael Hardy
Citations: 110
h-index: 5
Max Lamparth
Max Lamparth
Stanford University
Citations: 986
h-index: 12
Kevin Klyman
Kevin Klyman
Citations: 883
h-index: 15
Aarush Sinha
Aarush Sinha
Citations: 15
h-index: 3
N. Heath
N. Heath
Citations: 0
h-index: 0
Shalaleh Rismani
Shalaleh Rismani
Citations: 866
h-index: 11
Subramanyam Sahoo
Subramanyam Sahoo
Citations: 6
h-index: 2
Michael A. Riegler
Michael A. Riegler
Citations: 16
h-index: 2
Wm. Matthew Kennedy
Wm. Matthew Kennedy
Citations: 13
h-index: 1
Andrew Tran
Andrew Tran
Citations: 44
h-index: 3
A. Kornilova
A. Kornilova
Citations: 251
h-index: 5
Damian Stachura
Damian Stachura
Citations: 42
h-index: 2
F. Friedrich
F. Friedrich
Citations: 22,123
h-index: 60
Anoop Mishra
Anoop Mishra
Citations: 71
h-index: 5
Yixiong Hao
Yixiong Hao
Georgia Tech
Citations: 9
h-index: 2
Andreas Loehr
Andreas Loehr
Citations: 1
h-index: 1
Ruchira Dhar
Ruchira Dhar
Citations: 34
h-index: 4
Sree Harsha Nelaturu
Sree Harsha Nelaturu
Citations: 79
h-index: 3
Drishti Sharma
Drishti Sharma
Citations: 13
h-index: 2
I. Khire
I. Khire
Citations: 8
h-index: 1
Amit Saha
Amit Saha
Citations: 1
h-index: 1
Kabir Manghnani
Kabir Manghnani
Citations: 129
h-index: 3
M. Lin
M. Lin
Citations: 76
h-index: 2
Yanan Jiang
Yanan Jiang
Citations: 117
h-index: 6
Yilin Huang
Yilin Huang
Citations: 7
h-index: 1
Jessica Ji
Jessica Ji
Citations: 19
h-index: 3
A. Hofmann
A. Hofmann
Citations: 0
h-index: 0
Mubashara Akhtar
Mubashara Akhtar
King's College London
Citations: 217
h-index: 5
Nuno Moniz
Nuno Moniz
Citations: 0
h-index: 0
Yacine Jernite
Yacine Jernite
Hugging Face
Citations: 11,816
h-index: 28
Zeerak Ta-lat
Zeerak Ta-lat
Citations: 54
h-index: 1

AI 평가 결과는 대량으로 생성되지만, 리더보드, 모델 카드, 벤치마크 논문 및 회사 블로그에 따라 일관성이 없습니다. 이로 인해 독자는 다양한 출처의 결과를 신뢰성 있게 비교하거나, 보고서에서 누락된 내용을 파악하거나, 종합적인 주장을 해당 근거 자료와 연결하기 어렵습니다. 최근 노력들은 개별 구성 요소를 다루지만, 세 가지 격차를 남깁니다. 즉, 평가 수명 주기 전체를 포괄하지 않고 단일화된 해석 가능한 기록으로 구성되지 않으며, 동일한 증거에 대해 다양한 이해 관계자가 제시하는 질문의 차이를 구분하지 않는 정적인 표현을 지정하며, 실제 적용을 위한 추출 인프라가 부족하여 실증적으로 검증되지 않은 제안 단계에 머물러 있습니다. 본 연구에서는 벤치마크 메타데이터, 평가 실행 데이터 및 모델 메타데이터를 통합된 기록으로 구성하는 운영 보고 레이어인 exttt{EvalCards}를 제시합니다. (1) 52편의 논문과 10건의 이해 관계자 인터뷰를 통해 체계적으로 도출한 보고 스키마를 기반으로 하고, (2) 연구 및 비연구 대상 독자를 위한 맞춤형 모드를 제공하여 재현성, 문서 완성도, 출처 및 위험 요소, 점수 비교 가능성을 나타내는 네 가지 해석 신호를 구현하며, (3) exttt{EvalCards}를 5,816개의 모델, 635개의 벤치마크 및 101,843개의 결과에 적용하는 모니터링 도구를 배포하여 현재 보고 방식의 체계적인 문제점을 파악합니다.

Original Abstract

AI evaluation results are produced at scale but reported inconsistently across leaderboards, model cards, benchmark papers, and company blogs. The cost is interpretive: readers cannot reliably compare results across sources, identify what a report omits, or trace an aggregate claim to its underlying evidence. Recent efforts address isolated components but leave three gaps: they cover only narrow slices of the evaluation lifecycle and do not compose into a single interpretable record; they specify static representations that do not differentiate the questions different stakeholders bring to the same evidence; and they remain proposals on paper, lacking the extraction infrastructure required for adoption at scale. We present \EvalCards{}, an operational reporting layer that composes benchmark metadata, evaluation run data, and model metadata into a unified record. We (1) derive a reporting schema from a structured review of 52 papers and 10 stakeholder interviews, (2) implement four interpretive signals (reproducibility, documentation completeness, provenance and risk, and score comparability), rendered through reader modes calibrated to research and non-research audiences, and (3) deploy a monitoring tool that applies \EvalCards{} across 5,816 models, 635 benchmarks, and 101,843 results, surfacing systematic gaps in current reporting practice.

0 Citations
0 Influential
30 Altmetric
150.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!