2607.00304v1 Jul 01, 2026 cs.LG

평가 시스템의 한계 탐색: 평가자와 에이전트 간 조건 하에서의 편향-신뢰성 균형에 대한 실증적 연구 (11가지 조건)

Mapping the Evaluation Frontier: An Empirical Survey of the Bias-Reliability Tradeoff Across Eleven Evaluator-Agent Conditions

Zewen Liu
Zewen Liu
Citations: 231
h-index: 7

편향-신뢰성 균형 이론은 LLM 평가 시스템이 평가자 결합(gamma), 전략 다양성(H) 및 소규모 샘플 측정 신뢰도(CV(N))의 세 가지 요소를 동시에 최적화할 수 없다는 것을 시사하며, 이는 고정된 샘플 크기 N에서 (gamma, H, CV) 공간에 의해 제약됩니다. 기존 연구는 단일 연구에서 완전한 지표를 갖춘 5가지 조건에 기반하고 있습니다. 본 연구에서는 실증적 근거를 확장하여 11가지 조건으로 분석을 진행했으며, 모든 11가지 조건에서 gamma와 H 값을 측정하고, 충분한 시드(N >= 5)를 가진 7가지 조건에서 CV(N=5) 값을 측정했습니다. 5가지 조건에서는 (gamma, H, CV)의 세 가지 지표를 모두 제공합니다. 분석 결과는 균형 관계를 확인시켜줍니다. 낮은 평가자 결합(gamma < 0.2)을 보이는 조건은 높은 측정 노이즈(CV(N=5) > 1.0)를 나타내는 반면, 강한 결합(gamma > 0.9)을 보이는 조건은 낮은 노이즈(CV(N=5) < 0.16)를 달성합니다. 전략 다양성과 평가자 결합 간의 상관관계 r(H, gamma) = -0.989 (n=5, GPT-4o 조건 제외)는 평가자 결합이 전략 다양성을 저해함을 확인시켜줍니다. 4가지 GPT-4o 조건에서는 모든 시드에서 gamma가 0.000이고 H가 1.000인 패턴을 보였으며, 이는 2026년 6월 GPT-4o API의 버전 변경으로 인한 현상으로 추정됩니다. 어떤 조건에서도 {gamma < 0.2, CV(N=5) < 0.3} 영역에 속하지 않습니다. 본 연구에서 얻은 모든 조건별 지표는 평가자 비교를 위한 표준화된 벤치마크 데이터 세트로 공개합니다.

Original Abstract

The bias-reliability tradeoff conjectures that LLM evaluation systems are constrained in (gamma, H, CV) space, where evaluator coupling (gamma), strategy diversity (H), and small-sample measurement reliability (CV(N)) cannot be simultaneously optimized at fixed sample size N. Prior evidence rests on n=5 conditions with complete metrics from a single study. We expand the empirical base to 11 conditions, measuring gamma and H for all 11 (nine with valid weight vectors) and CV(N=5) for seven with sufficient seeds (N >= 5). Five conditions provide the complete (gamma, H, CV) triple. The data confirm the trade-off: conditions with low evaluator coupling (gamma < 0.2) exhibit high measurement noise (CV(N=5) > 1.0), while conditions with strong coupling (gamma > 0.9) achieve low noise (CV(N=5) < 0.16). The correlation r(H, gamma) = -0.989 (n=5, excluding GPT-4o conditions) confirms that evaluator coupling suppresses strategy diversity. Four GPT-4o conditions show gamma=0.000 and H=1.000 across all seeds -- a pattern we attribute to version drift in the June 2026 GPT-4o API. No condition occupies the region {gamma < 0.2, CV(N=5) < 0.3}. We release all per-condition metrics as a standardized benchmark dataset for evaluator comparison.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!