인공지능 평가에 아이템 반응 이론을 신뢰할 수 있는가?
Can We Trust Item Response Theory for AI Evaluation?
인공지능 성능 평가 지표는 모델의 능력 추정, 시스템 순위 결정, 유용한 예시 선택 및 지표 품질 진단을 위해 점차적으로 개별 항목 수준의 통계 모델, 특히 아이템 반응 이론(IRT)을 활용하고 있습니다. 그러나 인공지능 성능 평가 데이터는 종종 인간 시험 데이터의 체계를 벗어납니다. 이는 기존 IRT 추정 도구가 원래 개발된 환경과 다르기 때문입니다. 일반적으로 지표는 평가되는 모델 수가 적고, 항목 수는 훨씬 많으며, 능력 분포가 왜곡되거나, 클러스터링되어 있거나, 다중 모드를 가질 수 있습니다. 본 연구에서는 이러한 체계 불일치가 인공지능 평가를 위한 IRT 모델링의 신뢰성에 어떤 영향을 미치는지 분석합니다. 널리 사용되는 6개의 LLM 지표에서 파생된 항목 매개변수 및 능력 분포를 사용하여 세 가지 일반적인 IRT 모델 하에서 응답 행렬을 시뮬레이션하고, 최근 지표 연구에 사용된 네 가지 추정 도구(경계 최대 우도법, 마르코프 체인 몬테카를로 방법, 변분 추론, 신경망 기반 유사 Siamese 추정기)를 비교합니다. 18,000개의 시뮬레이션 조건을 통해 모델 순위, 예측 성능 및 항목 특성에 대한 IRT 추론의 신뢰성뿐만 아니라 계산 가능성과 확장성을 체계적으로 평가합니다. 결과는 기존 추정기가 대규모 지표 환경에서 실행 불가능해질 수 있으며, 확장 가능한 추정기는 작은 규모 또는 비정상 분포를 가진 모델 집합에서 부정확한 항목 수준 및 순위 추론을 초래할 수 있음을 보여줍니다. 본 연구는 잠재적 특성 모델이 인공지능 성능 평가 주장을 안정적으로 뒷받침하는지 또는 왜곡 위험이 있는지, 그리고 신뢰할 수 있는 사용을 위해 필요한 표본 크기와 진단 기준은 무엇인지에 대한 정보를 제공합니다.
AI benchmarks increasingly leverage item-level statistical models, particularly item response theory (IRT), to estimate model capabilities, rank systems, select informative examples, and diagnose benchmark quality. However, AI benchmark data often departs from the data regime of human testing, for which standard IRT estimation tools were originally developed: benchmarks typically involve fewer evaluated models, far more items, and capability distributions that may be skewed, clustered, or multimodal. We examine how these regime mismatches challenge the reliability of IRT modeling for AI evaluation. Using item parameters and capability distributions derived from six widely used LLM benchmarks, we simulate response matrices under three common IRT models and compare four estimation tools used in recent benchmark studies: marginal maximum likelihood, Markov chain Monte Carlo, variational inference, and a neural pseudo-Siamese estimator. Across 18,000 simulation conditions, we systematically evaluate computational feasibility, scalability, and the reliability of IRT inferences about model rankings, predicted performance, and item characteristics. Results show that classical estimators can become infeasible in large benchmark settings, whereas scalable estimators can produce unreliable item-level and ranking inferences with small or nonnormally distributed model sets. This study identifies when latent trait models reliably support or risk distorting AI benchmarking claims, and what sample sizes and diagnostics are needed for trustworthy use.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.