모든 오답이 중요하다: LLM 객관식 벤치마크를 위한 옵션 레벨 심리 측정 방법
Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks
대부분의 객관식 질문(MCQ) 벤치마크는 대규모 언어 모델(LLM)을 평가할 때 정답 여부만으로 판단합니다. 이러한 이분법적 점수는 모든 오답을 동일하게 취급하지만, LLM이 부정확한 선택지 중에서 보이는 선호도는 LLM의 행동과 능력에 대한 체계적인 정보를 담고 있을 수 있습니다. 본 연구에서는 LLM-NRM(LLM Nominal Response Model)이라는 옵션 기반 심리 측정 프레임워크를 소개합니다. LLM-NRM은 모든 응답 선택지에 대한 분포를 모델링하여 LLM의 능력을 추정하고, 동시에 각 선택지의 특성을 파악하며, 모델별 응답 보정 정확도, 위치 선호도, 난이도에 따른 대체 행동 패턴 등을 구분합니다. 189개의 LLM과 14개 벤치마크에서 총 31,554개의 항목을 사용하여 LLM-NRM은 이분법적 모델 및 기존의 기본적인 방법보다 더 정확하게 LLM-항목 간 상호작용을 예측하며, 추정된 능력 값은 외부 인간 선호도 평가 시스템인 Arena.ai Elo 순위와 0.920이라는 가장 높은 Spearman 상관관계를 보입니다. 정답 여부 외에 오답 선택지의 정보는 항목당 Fisher Information을 +101% 증가시키며, 오답 데이터만으로도 LLM의 능력을 완전히 추정할 수 있으며, 이때 Spearman 상관계수는 0.943입니다. 학습된 항목 파라미터는 효율적인 벤치마킹에도 활용될 수 있는데, 41개의 선택된 항목만으로는 전체 항목 집합의 순위를 0.85의 Kendall 상관관계로 유지할 수 있으며, 이는 데이터 양을 약 770배 줄이는 효과를 가져옵니다. 결론적으로, 우리는 오답이 단순한 실수가 아니라, LLM의 능력을 측정하는 데 중요한 정보를 담고 있다는 것을 보여줍니다.
Most multiple-choice question (MCQ) benchmarks evaluate Large Language Models (LLMs) only by whether they select the correct answers. This binary scoring treats all incorrect responses alike, even though an LLM's preferences among incorrect options may contain systematic and useful information about its behavior and ability. We introduce the LLM Nominal Response Model (LLM-NRM), an option-aware psychometric framework that models the full distribution over answer choices to jointly estimate LLM ability and option-level item characteristics, while separating model-specific response calibration sharpness, positional preference, and difficulty-dependent fallback behavior. Across 189 LLMs and 31,554 items from 14 benchmarks, LLM-NRM predicts held-out LLM-item interactions more accurately than binary Item Response models and conventional nominal-response baselines, and its ability estimates achieve the strongest Spearman correlation of 0.920 with the external human-preference Arena.ai Elo leaderboard. Distractor identity contributes +101% additional Fisher Information per item beyond correctness, and incorrect responses alone recover full-information ability estimates with Spearman 0.943. The learned item parameters also enable efficient benchmarking, where 41 selected items preserve the full-bank ranking with Kendall's correlation 0.85, corresponding to a 770 times reduction. In conclusion, we show that incorrect answers carry distinct and useful measurement information rather than representing equivalent mistakes.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.