2604.06996v1 Apr 08, 2026 cs.CL

러브릭 기반 대규모 언어 모델 평가 시 나타나는 자기 선호 편향

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models

José Pombal
José Pombal
Sword Health
Citations: 751
h-index: 10
Ricardo Rei
Ricardo Rei
Citations: 430
h-index: 11
Andr'e F. T. Martins
Andr'e F. T. Martins
Citations: 143
h-index: 5

LLM-as-a-judge 방식은 대규모 언어 모델(LLM)의 출력 결과를 평가하는 사실상 표준적인 방법이 되었습니다. 그러나 평가자는 자기 선호 편향(Self-Preference Bias, SPB)을 보이는 경향이 있는데, 이는 자신이 생성한 출력 또는 자신이 속한 모델 패밀리의 출력에 대해 더 높은 점수를 부여하는 현상을 의미합니다. 이러한 편향은 평가 결과를 왜곡하고, 특히 자체 개선을 반복하는 환경에서 모델 개발을 저해합니다. 본 연구는 러브릭 기반 평가에서 발생하는 SPB 현상에 대한 최초의 연구입니다. 러브릭 기반 평가는 평가자가 개별 평가 기준에 대해 이분법적인 판단을 내리는, 점점 더 보편화되는 벤치마킹 방식입니다. IFEval이라는, 프로그램적으로 검증 가능한 러브릭을 사용하는 벤치마크를 사용하여, 평가 기준이 완전히 객관적일 때에도 SPB가 지속된다는 것을 확인했습니다. 생성 모델이 특정 기준에 부합하지 않는 경우, 평가자는 생성된 출력이 자신의 것이라면 최대 50% 더 높은 확률로 이를 만족하는 것으로 잘못 판단하는 경향이 있습니다. 또한, 다른 평가 방식과 마찬가지로 여러 평가자를 조합하는 것이 SPB를 완화하는 데 도움이 되지만, 완전히 제거하지는 못한다는 것을 확인했습니다. HealthBench라는 의료 챗봇 벤치마크에서, 주관적인 러브릭을 사용하는 경우 SPB가 모델 점수를 최대 10점까지 왜곡할 수 있는데, 이는 최첨단 모델을 순위 매할 때 중요한 차이를 만들 수 있는 수준입니다. 본 연구에서는 이러한 환경에서 SPB를 유발하는 요인을 분석한 결과, 부정적인 러브릭, 극단적인 러브릭 길이, 그리고 응급 환자 이송과 같은 주관적인 주제가 특히 SPB에 취약한 것으로 나타났습니다.

Original Abstract

LLM-as-a-judge has become the de facto approach for evaluating LLM outputs. However, judges are known to exhibit self-preference bias (SPB): they tend to favor outputs produced by themselves or by models from their own family. This skews evaluations and, thus, hinders model development, especially in settings of recursive self-improvement. We present the first study of SPB in rubric-based evaluation, an increasingly popular benchmarking paradigm where judges issue binary verdicts on individual evaluation criteria, instead of assigning holistic scores or rankings. Using IFEval, a benchmark with programmatically verifiable rubrics, we show that SPB persists even when evaluation criteria are entirely objective: among rubrics where generators fail, judges can be up to 50\% more likely to incorrectly mark them as satisfied when the output is their own. We also find that, similarly to other evaluation paradigms, ensembling multiple judges helps mitigate SPB, but without fully eliminating it. On HealthBench, a medical chat benchmark with subjective rubrics, we observe that SPB skews model scores by up to 10 points, a potentially decisive margin when ranking frontier models. We analyze the factors that drive SPB in this setting, finding that negative rubrics, extreme rubric lengths, and subjective topics like emergency referrals are particularly susceptible.

5 Citations
1 Influential
5.5 Altmetric
34.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!