2605.27914v1 May 27, 2026 cs.CL

결과가 말하게 하라: LLM 행동 평가를 위한 재현성 중심 패러다임

Let the Results Speak: A Replication-First Paradigm for LLM Behavioral Benchmarking

Junchen Wan
Junchen Wan
Citations: 91
h-index: 4
Yaopei Liu
Yaopei Liu
Citations: 121
h-index: 2
Yumin Huang
Yumin Huang
Citations: 21
h-index: 2
Lei Wang
Lei Wang
Citations: 78
h-index: 3
Peng Ding
Peng Ding
Citations: 10
h-index: 2

LLM의 행동, 즉 공감 능력, 자제력, 조절된 감정 표현 등에 대한 주관적인 평가는 어렵습니다. 인간 간의 일치도는 이러한 특성에 대해 약 0.45 수준으로 제한적이며, LLM을 판단자로 사용하는 것은 순환 오류를 초래할 위험이 있습니다. 대상 모델의 학습 데이터와 동일한 데이터를 사용한 판단자는 독립적으로 검증할 수 없습니다. 단일 인간 평가자의 합의에 의존하는 방식은 인간들조차 의견이 다른 능력 영역에는 적용될 수 없습니다. 저희는 재현성을 최우선으로 하는 패러다임을 제안합니다. 이는 하나의 평가 그룹에 의존하는 대신, 신뢰성(K번 실행), 아키텍처적으로 다양한 판단자 간의 교차 검증, 이전 학습 데이터셋 출신 판단자를 통한 역사적 보정, 그리고 사전 등록된 예측이라는 네 가지 직교적인 특성을 통해 평가 도구를 검증합니다. 저희는 이 패러다임을 감정적 지원 분야에 적용하여, 평가 기준을 반복적으로 개선하며 데이터를 기반으로 자체적으로 발전시켰습니다. 즉, 사전에 정의된 차원이 아닌 9가지 차원으로 구성됩니다. 사전 등록은 10개의 반증 가능한 가설과 11개의 예측을 포함하며, 이는 테스트 데이터 수집 전에 결정되었습니다. 이 패러다임을 8개 계열의 49개 모델에 적용한 결과, 집계 점수가 숨기는 문제점을 드러낼 수 있었습니다. 예를 들어, 공감적인 맥락에서 모델이 원치 않는 해결책을 제시하지 않는 정도인 '조언-자제력' 측면에서, gpt-5는 gpt-4.1보다 1.87점이 낮았고, Opus-4.7은 Opus-4.6보다 0.629점이 낮았습니다. 이때 집계 점수는 변화가 없었습니다. 이 회귀 분석은 세 번의 사용자 대리 평가 변경에도 불구하고 95% 수준으로 유지되었으며, 5개 계열의 판단자 그룹과 17개월의 시간 간격을 두고도 재현되었습니다. 또한, 74개의 실제 ESConv 대화 데이터에 대해서도 유효성이 검증되었으며(rho는 [0.749, 0.850] 범위), 평가 도구의 순위 상관 계수(Krippendorff alpha)는 0.91로 매우 높았습니다. 이 패러다임은 부가적으로 평가 기준 개선으로 극복할 수 있는 평가 기준의 한계와, 시나리오 또는 데이터셋 변경이 필요한 구조적인 한계를 구별하는 데 도움이 됩니다.

Original Abstract

Subjective evaluation of LLM behavior -- empathy, restraint, calibrated emotional tone -- is hard. Human inter-rater agreement on such qualities saturates near rho ~ 0.45, and an LLM-as-judge proxy alone risks circularity: a judge sharing the target's training cohort cannot independently verify it. Anchoring validity to a single human-rater consensus does not extend to capabilities where humans themselves disagree. We propose a replication-first paradigm: instead of anchoring on one rater group, we certify the instrument via four orthogonal properties -- reliability across K runs, cross-instrument replication across architecturally distinct judges, historical-footprint calibration via judges from earlier training cohorts, and pre-registered prediction. We test it on emotional accompaniment by letting the rubric self-evolve data-driven across iterations: the dimensions are not pre-stipulated and the procedure stabilizes to a 9-dimension set. Pre-registration applies to 10 falsifiable hypotheses and 11 forward predictions, committed before any test data was collected. Applied to 49 models across 8 families, the paradigm surfaces what aggregate scores hide. On advice-restraint -- whether a model refrains from giving unsolicited solutions in empathic contexts -- gpt-5 falls 1.87 points from gpt-4.1 and Opus-4.7 falls 0.629 from Opus-4.6, while aggregate scores stay flat. The regression survives three user-proxy swaps (95% of magnitude), replicates across a 5-family judge stack and a 17-month cohort gap, and persists on 74 held-out real ESConv conversations (rho in [0.749, 0.850]); the instrument reaches ordinal Krippendorff alpha = 0.91. As a by-product, the paradigm acts as a saturation-source diagnostic, separating instrumental ceilings (breakable by rubric refinement) from structural ceilings (needing scenario or roster intervention).

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!