CalibratedRubric: 개방형 LLM 평가를 위한 작업 적응형 루브릭 은행
CalibratedRubric: Task-Adaptive Rubric Banks for Open-Ended LLM Evaluation
개방형 LLM 출력의 신뢰성 있는 평가는 세밀한 기준(rubric)을 필요로 하지만, 전문가의 직접적인 기준 개발은 비용이 많이 들고 확장하기 어렵습니다. 기존 자동화된 시스템은 엄격한 평가자 합치도와 이분법적 변동 필터를 사용하는데, 이는 측정 가능한 기준과 유용한 정보를 제공하는 기준을 구별할 수 없습니다. 본 연구에서는 작업에 적응형으로 조정되는 프레임워크인 CalibratedRubric을 소개합니다. CalibratedRubric은 유형별 점수 부여, 베이지안 루브릭-측정 가능성 필터링, 그리고 아이템 반응 이론(IRT) 기반의 은행 조립 방식을 결합합니다. CalibratedRubric은 Beta--Bernoulli 합치도 사후 추정을 통해 각 기준의 측정 가능성을 추정하고, 관찰된 성능 범위에 대한 간결한 루브릭 은행을 구축하기 위해 부분적으로 가산적인 정보-커버리지 목표를 사용합니다. 금융, 의료, 일반 및 법률 벤치마크에서 측정 가능성 필터링은 JudgmentBench에서의 인간-참조 기준 일치도를 κ=0.604에서 0.743으로 향상시켰습니다. IRT 기반의 탐욕적 선택 방식은 모든 여섯 가지 평가된 응답 블록에서 무작위 선택보다 더 높은 순위 재현성을 제공하며, FinResearchBench 의사 결정 지원 작업에서 목표 상관 관계를 달성하는 데 131개 대신 49개의 루브릭만 필요합니다. 작업 레이블의 미세 조정은 시스템 간 분산을 더욱 줄여주며, 이는 작업 적응형 점수 부여 방식의 실제적인 유용성을 뒷받침합니다. 이러한 결과는 CalibratedRubric이 충분한 평가자 중복을 통해 교정 효과를 얻을 수 있는 효율적이고 불확실성에 대한 인식을 갖춘 개방형 LLM 평가 방법임을 시사합니다.
Reliable evaluation of open-ended LLM outputs requires fine-grained rubrics, yet expert curation is costly and difficult to scale. Existing automated pipelines rely on strict judge unanimity and binary variance filters, which cannot distinguish measurable rubrics from informative ones. We introduce CalibratedRubric, a task-adaptive framework that combines type-specific scoring, Bayesian rubric-measurability filtering, and item response theory (IRT)-based bank assembly. CalibratedRubric estimates each rubric's measurability with a Beta--Bernoulli agreement posterior and uses a submodular information-coverage objective to construct compact rubric banks over the observed capability range. Across financial, healthcare, general, and legal benchmarks, measurability filtering improves human-gold agreement on JudgmentBench from $κ=0.604$ to $0.743$. IRT-based greedy selection improves cross-fitted rank fidelity over random selection across all six evaluated response blocks and requires only 49 rather than 131 rubrics to reach the target correlation on FinResearchBench decision-support tasks. Task-label perturbations further reduce system separation, confirming the practical relevance of task-adaptive scoring. These results support CalibratedRubric as an efficient, uncertainty-aware approach to open-ended LLM evaluation, with calibration gains depending on sufficient judge redundancy.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.