평가 시스템의 한계 탐색: 평가자와 에이전트 간 조건 하에서의 편향-신뢰성 균형에 대한 실증적 연구 (11가지 조건)
Mapping the Evaluation Frontier: An Empirical Survey of the Bias-Reliability Tradeoff Across Eleven Evaluator-Agent Conditions
편향-신뢰성 균형 이론은 LLM 평가 시스템이 평가자 결합(gamma), 전략 다양성(H) 및 소규모 샘플 측정 신뢰도(CV(N))의 세 가지 요소를 동시에 최적화할 수 없다는 것을 시사하며, 이는 고정된 샘플 크기 N에서 (gamma, H, CV) 공간에 의해 제약됩니다. 기존 연구는 단일 연구에서 완전한 지표를 갖춘 5가지 조건에 기반하고 있습니다. 본 연구에서는 실증적 근거를 확장하여 11가지 조건으로 분석을 진행했으며, 모든 11가지 조건에서 gamma와 H 값을 측정하고, 충분한 시드(N >= 5)를 가진 7가지 조건에서 CV(N=5) 값을 측정했습니다. 5가지 조건에서는 (gamma, H, CV)의 세 가지 지표를 모두 제공합니다. 분석 결과는 균형 관계를 확인시켜줍니다. 낮은 평가자 결합(gamma < 0.2)을 보이는 조건은 높은 측정 노이즈(CV(N=5) > 1.0)를 나타내는 반면, 강한 결합(gamma > 0.9)을 보이는 조건은 낮은 노이즈(CV(N=5) < 0.16)를 달성합니다. 전략 다양성과 평가자 결합 간의 상관관계 r(H, gamma) = -0.989 (n=5, GPT-4o 조건 제외)는 평가자 결합이 전략 다양성을 저해함을 확인시켜줍니다. 4가지 GPT-4o 조건에서는 모든 시드에서 gamma가 0.000이고 H가 1.000인 패턴을 보였으며, 이는 2026년 6월 GPT-4o API의 버전 변경으로 인한 현상으로 추정됩니다. 어떤 조건에서도 {gamma < 0.2, CV(N=5) < 0.3} 영역에 속하지 않습니다. 본 연구에서 얻은 모든 조건별 지표는 평가자 비교를 위한 표준화된 벤치마크 데이터 세트로 공개합니다.
The bias-reliability tradeoff conjectures that LLM evaluation systems are constrained in (gamma, H, CV) space, where evaluator coupling (gamma), strategy diversity (H), and small-sample measurement reliability (CV(N)) cannot be simultaneously optimized at fixed sample size N. Prior evidence rests on n=5 conditions with complete metrics from a single study. We expand the empirical base to 11 conditions, measuring gamma and H for all 11 (nine with valid weight vectors) and CV(N=5) for seven with sufficient seeds (N >= 5). Five conditions provide the complete (gamma, H, CV) triple. The data confirm the trade-off: conditions with low evaluator coupling (gamma < 0.2) exhibit high measurement noise (CV(N=5) > 1.0), while conditions with strong coupling (gamma > 0.9) achieve low noise (CV(N=5) < 0.16). The correlation r(H, gamma) = -0.989 (n=5, excluding GPT-4o conditions) confirms that evaluator coupling suppresses strategy diversity. Four GPT-4o conditions show gamma=0.000 and H=1.000 across all seeds -- a pattern we attribute to version drift in the June 2026 GPT-4o API. No condition occupies the region {gamma < 0.2, CV(N=5) < 0.3}. We release all per-condition metrics as a standardized benchmark dataset for evaluator comparison.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.