2606.19714v1 Jun 18, 2026 stat.ML

AURA: LLM 기반 평가 시스템 감사 및 개선을 위한 불확실성 인지적 적응형 정제 방법

AURA: Adaptive Uncertainty-aware Refinement for LLM-as-a-Judge Auditing

C. Yeh
C. Yeh
Citations: 6,980
h-index: 6
Lei Ding
Lei Ding
Citations: 33
h-index: 3
Zilong Zhang
Zilong Zhang
Citations: 0
h-index: 0
Yi-Ting Hung
Yi-Ting Hung
Citations: 4
h-index: 1
Weiyi He
Weiyi He
Citations: 0
h-index: 0
Junxi Zhang
Junxi Zhang
Citations: 50
h-index: 2

대규모 언어 모델(LLM)은 대규모 인간 평가가 비용이 많이 들고 확장하기 어려우므로, 개방형 생성 결과에 대한 판단자로 점점 더 많이 사용되고 있습니다. 하지만 LLM의 선호도는 여전히 인간의 판단을 완벽하게 반영하지 못합니다. 기존 감사 파이프라인은 종종 인간 주석, 휴리스틱 필터링 또는 강력한 평가 시스템의 출력과 같은 사전 정의된 신뢰할 수 있는 예제 집합이나 깨끗한 감독 신호가 존재한다고 가정합니다. LLM 평가에서는 이러한 가정이 취약합니다. 초기 데이터 분할은 평가자의 편향을 상속할 수 있으며, 인간 검증은 일반적으로 안정적인 그룹을 대규모로 정의하기에는 부족합니다. 본 논문에서는 선택된 인간 검증을 활용하여 LLM 기반 평가 시스템의 판단 결과에 대한 감사를 위한 적응형 불확실성 인지적 정제 프레임워크인 AURA를 제안합니다. AURA는 반복적으로 인간 일관성 신호를 학습하고, 신뢰할 수 있는 증거를 전파하며, 인간 검토를 위해 불확실성이 높은 비교 항목을 우선순위로 지정합니다. 핵심 아이디어는 평가 시스템에 대한 신뢰도를 잠재적인 값으로 취급하고, 증거가 축적됨에 따라 점진적으로 정제하는 것입니다. 우리는 간결한 공식화, 안정적인 정제 절차 및 합성 데이터와 실제 LLM 답변 데이터를 모두 사용한 종합적인 평가 결과를 제시합니다.

Original Abstract

Large language models (LLMs) are increasingly used as judges for open-ended generation, as large-scale human evaluation is often expensive and difficult to scale, yet their preferences remain imperfect proxies for human judgment. Existing auditing pipelines often assume that a reliable subset of examples or clean supervision signals are available beforehand, for example from human annotation, heuristic filtering, or the outputs of strong judges. In LLM evaluation, this assumption is fragile: the initial split may inherit judge bias, while human verification is typically too scarce to define stable groups at scale. We propose AURA, an adaptive uncertainty--aware refinement framework for auditing pairwise LLM--as--a--judge decisions under selected human verification. AURA iteratively learns a human-consistency signal, propagates reliable evidence, and prioritizes uncertain comparisons for human review. The key idea is to treat trust in a judge as a latent quantity that is progressively refined as evidence accumulates. We provide a compact formulation, a stable refinement procedure, and a comprehensive evaluation on both synthetic and real pairwise LLM-answer data.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!