2606.29784v1 Jun 29, 2026 stat.ME

HERO: 과거 데이터를 활용하여 생성 모델 평가의 신뢰성과 민감성을 향상시키는 방법

HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data

Xinrui Ruan
Xinrui Ruan
Citations: 2
h-index: 1
Zhenyu Zhao
Zhenyu Zhao
Citations: 1
h-index: 1
Waverly Wei
Waverly Wei
Citations: 70
h-index: 6
Yueshan Zhang
Yueshan Zhang
Citations: 0
h-index: 0
Sui Huang
Sui Huang
Citations: 124
h-index: 5
Jingshen Wang
Jingshen Wang
Citations: 59
h-index: 5
Zeyu Zheng
Zeyu Zheng
Citations: 5,358
h-index: 20

신뢰성 있는 생성 AI 모델은 출력 품질을 평가하기 위해 전문가의 인간 주석에 크게 의존하지만, 이러한 "골드" 레이블을 수집하는 데는 비용이 많이 들고 양도 제한적입니다. 따라서 조직은 종종 골드 레이블의 대안으로 크라우드소싱 작업자 또는 벤더 어노테이터로부터 얻은 방대한 데이터인 "실버" 레이블을 수집합니다. 그러나 실제 평가 대상은 여전히 골드 레이블이므로, 노이즈가 많은 실버 레이블을 무분별하게 결합하면 편향이 발생할 수 있으며, 희소한 골드 레이블을 기반으로 구축된 추정치는 모델 성능 격차를 해소하기 위해 높은 분산을 가질 수 있습니다. 모델 평가는 일회성 작업이 아니라 지속적인 운영 활동이 되었으며, 평가 라운드는 모델 버전, 릴리스 및 콘텐츠 도메인 전반에 걸쳐 반복됩니다. 자연스러운 질문은 이전의 과거 평가 데이터를 사용하여 새로운 평가 라운드를 개선할 수 있는지 여부입니다. 본 논문에서는 과거 데이터를 활용하여 모델 성능 평가에서 편향을 줄이고 (신뢰성 향상) 분산을 낮추는 (민감성 향상) 새로운 프레임워크인 HERO (History Enhanced RObust model evaluation)를 소개합니다. HERO는 과거 골드 어노테이션으로부터 학습된 실버 레이블러의 성능을 보정하고, 과거 데이터에 정확하게 측정된 공변량 정보를 활용하여 추정기를 안정화합니다. HERO는 다양한 일반적인 평가 작업에 널리 적용될 수 있으며, 현재 라운드에 일부 과거 레이블러만 참여하는 경우에도 유효합니다. 본 논문에서는 편향 및 분산 감소 조건에 대한 분석을 제시하고, 시뮬레이션 연구를 통해 HERO의 성능을 입증하며, 실제 모델 평가 벤치마킹 데이터셋에서 HERO의 효과성을 보여줍니다.

Original Abstract

Reliable generative AI models critically rely on expert human annotations to evaluate output quality, yet these "gold" labels are expensive to collect and limited in quantity. Organizations thus often turn to collecting vast but noisy "silver" labels from crowdsourced workers or vendor annotators as proxies for gold labels. Because gold remains the evaluation target, naively aggregating noisy silver labels may introduce bias, and estimators built on sparsely observed gold labels may have high variance to resolve the model performance gaps that guide practical decisions. Model evaluation has become an ongoing operational practice rather than a one-time exercise, with evaluation rounds repeating across model versions, releases, and content domains. A natural question is whether the previous historical evaluation data can be used to improve each new round of evaluation. We introduce HERO (History Enhanced RObust model evaluation), a novel framework that uses historical data to suppress bias (improve reliability) and reduce variance (improve sensitivity) in model performance evaluation. HERO calibrates silver labelers' performance learned from historical gold annotations, and stabilizes the resulting estimator by anchoring it to covariate information measured with high precision in the historical data. HERO can be broadly applied across multiple common evaluation tasks, and remains valid when only a subset of historical labelers appears in the current round. We establish conditions under which the bias and variance reductions hold, showcase HERO's performance in simulation studies, and demonstrate its effectiveness on real-world model evaluation benchmarking datasets.

0 Citations
0 Influential
10 Altmetric
50.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!