2606.09409v1 Jun 08, 2026 cs.AI

더 정확한 것이 더 나아 보인다: 쌍대 비교를 통한 정확도 순위 분석

Correct Looks Better: Pairwise Comparisons Reveal Accuracy Rankings

Moritz Hardt
Moritz Hardt
Citations: 46
h-index: 3
Mina Remeli
Mina Remeli
Citations: 38
h-index: 3

쌍대 비교와 Elo 알고리즘과 같은 집계 방법은 생성 모델 평가에 중요한 역할을 하지만, 이러한 방식이 표면적인 스타일적 특징을 과장하거나 심사위원의 편향성을 반영할 수 있다는 우려가 존재합니다. 본 연구에서는 쌍대 비교를 통한 모델 순위가 실제 정확도를 기반으로 한 순위와 일치하는 경향이 강하다는 것을 확인했습니다. 5개의 대표적인 벤치마크를 자유 형식의 생성 평가 방식으로 변환하여 분석한 결과, Elo 알고리즘을 사용한 순위는 정확도 순위와 0.9 이상의 스피어만 상관관계를 보였으며, 심사위원의 역량이 부족할 경우 직접적인 평가 방식보다 훨씬 우수한 성능을 나타냈습니다. 또한, 스타일 및 심사위원 편향은 모델 순위에 미치는 영향이 미미하며, 이는 대부분의 판단이 두 후보 답변 모두 정확하거나 부정확한 쌍에 대해 이루어졌기 때문입니다. 이러한 쌍에 대한 분석 결과, 최종 답변 이후 반복되는 내용(echo)이 심사위원의 선호도에 중요한 영향을 미치는 것으로 나타났습니다.

Original Abstract

Pairwise comparisons combined with aggregation methods like Elo have become central to evaluating generative models, yet concerns remain that they reward superficial stylistic cues or display judge biases. In a more positive turn, we show that model rankings from pairwise comparisons strongly agree with ground-truth-based accuracy rankings when such ground truth is available for comparison. By converting five well-known benchmarks into free-form generative evaluations, we find that Elo rankings achieve a Spearman correlation above 0.9 with accuracy rankings and substantially outperform direct evaluation when the judge is weak. Furthermore, style and judge bias have only minor effects on model rankings, despite most judgments occurring on pairs where both candidate answers are correct (or incorrect). On such pairs, we find that repetition after the final answer (echo) is a causal driver of judge preference.

0 Citations
0 Influential
1.5 Altmetric
7.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!