평가 카드: AI 평가 보고를 위한 해석 레이어
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting
AI 평가 결과는 대량으로 생성되지만, 리더보드, 모델 카드, 벤치마크 논문 및 회사 블로그에 따라 일관성이 없습니다. 이로 인해 독자는 다양한 출처의 결과를 신뢰성 있게 비교하거나, 보고서에서 누락된 내용을 파악하거나, 종합적인 주장을 해당 근거 자료와 연결하기 어렵습니다. 최근 노력들은 개별 구성 요소를 다루지만, 세 가지 격차를 남깁니다. 즉, 평가 수명 주기 전체를 포괄하지 않고 단일화된 해석 가능한 기록으로 구성되지 않으며, 동일한 증거에 대해 다양한 이해 관계자가 제시하는 질문의 차이를 구분하지 않는 정적인 표현을 지정하며, 실제 적용을 위한 추출 인프라가 부족하여 실증적으로 검증되지 않은 제안 단계에 머물러 있습니다. 본 연구에서는 벤치마크 메타데이터, 평가 실행 데이터 및 모델 메타데이터를 통합된 기록으로 구성하는 운영 보고 레이어인 exttt{EvalCards}를 제시합니다. (1) 52편의 논문과 10건의 이해 관계자 인터뷰를 통해 체계적으로 도출한 보고 스키마를 기반으로 하고, (2) 연구 및 비연구 대상 독자를 위한 맞춤형 모드를 제공하여 재현성, 문서 완성도, 출처 및 위험 요소, 점수 비교 가능성을 나타내는 네 가지 해석 신호를 구현하며, (3) exttt{EvalCards}를 5,816개의 모델, 635개의 벤치마크 및 101,843개의 결과에 적용하는 모니터링 도구를 배포하여 현재 보고 방식의 체계적인 문제점을 파악합니다.
AI evaluation results are produced at scale but reported inconsistently across leaderboards, model cards, benchmark papers, and company blogs. The cost is interpretive: readers cannot reliably compare results across sources, identify what a report omits, or trace an aggregate claim to its underlying evidence. Recent efforts address isolated components but leave three gaps: they cover only narrow slices of the evaluation lifecycle and do not compose into a single interpretable record; they specify static representations that do not differentiate the questions different stakeholders bring to the same evidence; and they remain proposals on paper, lacking the extraction infrastructure required for adoption at scale. We present \EvalCards{}, an operational reporting layer that composes benchmark metadata, evaluation run data, and model metadata into a unified record. We (1) derive a reporting schema from a structured review of 52 papers and 10 stakeholder interviews, (2) implement four interpretive signals (reproducibility, documentation completeness, provenance and risk, and score comparability), rendered through reader modes calibrated to research and non-research audiences, and (3) deploy a monitoring tool that applies \EvalCards{} across 5,816 models, 635 benchmarks, and 101,843 results, surfacing systematic gaps in current reporting practice.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.