인지적 불확실성 평가: OOD 탐지 및 능동 학습을 넘어
Evaluating Epistemic Uncertainty: Beyond OOD Detection and Active Learning
현재 인지적 불확실성의 평가는 주로 데이터 분포 외부 탐지(OOD detection) 및 능동 학습과 같은 작업에 의존합니다. 그러나 이러한 작업에 대한 베이즈 최적의 의사 결정 전략은 일반적으로 인지적 불확실성을 정량화하는 데 사용되는 점수와 일치하지 않습니다. 우리는 인지적 거부 옵션 프레임워크를 기반으로, 인지적 불확실성이 줄일 수 있는 오류인 '후회(regret)'를 식별하는 능력에 따라 인지적 불확실성을 평가합니다. 선택적 예측을 커버리지, 예상 위험 및 후회를 고려한 제약 최적화 문제로 공식화하여, 최적의 선택자는 실제 값으로 주어진 확률론적(aleatoric) 및 인지적 불확실성의 임계값 기반 가중 합이라는 것을 증명했습니다. 이러한 이론적인 통일성은 최근의 불확실성 분리 연구의 약점을 드러냅니다. 우리는 학습된 구성 요소 간의 표준 상관 관계 지표가 반드시 실제 작동 유용성을 예측하지는 않는다는 것을 보여줍니다. 대신, 분해의 달성 가능한 위험, 후회 및 커버리지 표면을 평가하여 공동 분리와 유용성을 진단하는 방법을 제안합니다. 인간 주석이 풍부한 데이터 세트에서 표준 방법을 벤치마킹한 결과, 의사 결정 이론적 순위가 프록시 작업 순위와 크게 다를 수 있으며, 특정 기준에서 최상위를 차지하고 다른 기준에서는 최하위를 차지하는 방법 간에 페어별 순위 반전이 발생할 수 있습니다.
Current evaluation of epistemic uncertainty relies on tasks such as out-ofdistribution detection and active learning. However, the Bayes-optimal decision strategies for these tasks do not coincide with the scores commonly used to quantify epistemic uncertainty. Building on the epistemic reject-option framework, we evaluate epistemic uncertainty using its ability to identify regret, the reducible error. Formulating selective prediction as a constrained optimization over coverage, expected risk, and regret, we prove the optimal selector is a thresholded convex combination of the ground-truth aleatoric and epistemic uncertainties. This theoretical unification exposes a weakness in recent uncertainty disentanglement literature: we demonstrate that standard correlation metrics between learned components do not necessarily predict their actual operational utility. We instead propose to evaluate the achievable risk, regret, coverage surface of the decomposition as a diagnostic for joint disentanglement and utility. Benchmarking standard methods on datasets with dense human annotations reveals that decision-theoretic rankings can disagree substantially with proxy-task rankings, including pairwise rank inversions between methods that are top-ranked on one criterion and bottom-ranked on other.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.