불확실성 정량화의 의사 결정 연계 평가
Decision-Aligned Evaluation of Uncertainty Quantification
머신러닝에서 불확실성 추정은 일반적으로 음의 로그 우도(negative log-likelihood) 및 예상 교정 오차(expected calibration error)와 같은 일반적인 지표를 사용하여 평가됩니다. 그러나 이러한 지표에 대한 좋은 성능이 반드시 다운스트림 의사 결정에서의 높은 유용성을 의미하는 것은 아닙니다. 본 논문에서는 의사 결정 연계성(decision-alignment)이라는 기준을 제시하며, 이 기준은 어떤 평가 지표가 다운스트림 유용성과 의미 있게 연결되는지를 보여줍니다. 이 프레임워크를 적용하여, 널리 사용되는 많은 불확실성 지표들이 일반적인 의사 결정 문제와 일치하지 않거나, 다운스트림 작업에 대한 비정상적인 사전 가정을 내포하고 있음을 확인했습니다. 그런 다음, 본 논문에서는 사전 가중 유틸리티 지표(prior-weighted utility metrics)라는 특별한 형태의 적절한 스코어링 규칙을 제안합니다. 이 방법은 의사 결정 연계성을 갖는 불확실성 평가를 제공합니다. 벤치마크 실험과 실제 사례 연구에서, 제안하는 지표들은 기존 지표들과 달리 실제로 달성되는 의사 결정 유용성과 일관되게 연결되는 반면, 기존 지표들은 그렇지 않았습니다. 본 연구 결과는 현재의 불확실성 정량화 평가 프로토콜의 결점을 드러내며, 기존 지표를 의사 결정과 관련된 불확실성 평가로 확장하는 데 필요한 원칙적인 방법을 제시합니다.
Uncertainty estimates in machine learning are typically evaluated using generic metrics such as the negative log-likelihood and expected calibration error, yet good performance on such metrics does not necessarily imply high utility in downstream decisions. We introduce decision-alignment, a criterion that reveals which evaluation metrics meaningfully align with downstream utilities. Applying this framework, we show that many widely used uncertainty metrics are either misaligned with common decision problems or encode pathological prior beliefs about the downstream task. We then propose prior-weighted utility metrics, a special class of proper scoring rules that provides decision-aligned uncertainty evaluation. Across benchmark experiments and real-world case studies, our metrics consistently align with realized decision utility, while conventional metrics do not. Our results surface flaws in the current UQ evaluation protocol and offer a principled extension of existing metrics toward decision-relevant UQ evaluation.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.