2606.22792v1 Jun 22, 2026 cs.AI

확률성의 기원: 대규모 언어 모델의 불확실성 정량화를 위한 종합적 연구

The Origins of Stochasticity: Comprehensive Investigations on Uncertainty Quantification for Large Language Models

S. Liang
S. Liang
Citations: 64
h-index: 5
Xiang-Jun Ou
Xiang-Jun Ou
Citations: 0
h-index: 0
Xinmiao Hu
Xinmiao Hu
Citations: 0
h-index: 0
Rong Huang
Rong Huang
Citations: 1
h-index: 1
Jing Wang
Jing Wang
Citations: 9
h-index: 2
Shao-Qun Zhang
Shao-Qun Zhang
Citations: 2
h-index: 1

최근 대규모 언어 모델(LLM)의 발전은 정교한 추론과 콘텐츠 생성 기능을 가능하게 했지만, 내재된 확률성은 예측 신뢰성을 확보하는 데 상당한 어려움을 야기합니다. 전통적인 불확실성 분류 방식인 우연적 불확실성과 인식적 불확실성의 이분법은 개념적 기반을 제공하지만, LLM 생성의 다중 구성 요소 및 다단계 특성을 제대로 반영하지 못하며 다양한 불확실성 정량화(UQ) 방법의 효과를 평가하는 데 어려움을 겪습니다. 본 논문에서는 입력 수준, 파라미터 수준, 토큰 수준 및 디코딩 과정에서 발생하는 LLM의 불확실성을 체계적으로 분류하는 세분화된 불확실성 분류 방식을 제안합니다. 이에 따라 기존 UQ 방법을 베이지안 방법, 앙상블 방법, 합의 기반 방법 및 단일 패스 방식으로 분류했습니다. 또한 다양한 생성 환경과 지표를 포괄하는 종합적인 평가 프레임워크를 제시합니다. Qwen3, Llama 3.2, DeepSeek-V3 등 세 가지 주요 LLM 모델을 대상으로 TriviaQA, GSM8K, HumanEval 등의 벤치마크에서 21가지 일반적인 UQ 방법을 실증적으로 평가했습니다. 실험 결과는 다음과 같습니다 (i) UQ 방법의 효과는 작업 유형 및 생성 환경에 민감하며, (ii) Deg와 EigV와 같은 합의 기반 방법이 다른 UQ 접근 방식보다 일관되게 우수한 성능을 보이며, (iii) 모델 규모가 커질수록 불확실성 추정치가 낮아지는 경향이 있어 LLM 불확실성에 대한 경험적 스케일링 법칙을 제시합니다. 본 연구는 이론적 기원과 실제 적용 간의 격차를 해소하고, LLM 애플리케이션에서 불확실성을 체계적으로 정량화하기 위한 다재다능한 진단 도구를 제공합니다.

Original Abstract

Recent advancements in Large Language Models (LLMs) have enabled sophisticated reasoning and content generation, yet their inherent stochasticity poses significant challenges for ensuring predictive credibility. While traditional uncertainty taxonomy paradigms, such as the dichotomy of aleatoric and epistemic uncertainties, provide conceptual foundations, they often fail to capture the multi-component and multi-stage nature of LLM generation and struggle to evaluate the effectiveness of various Uncertainty Quantification (UQ) methods. In this paper, we propose a granular uncertainty taxonomy that systematically attributes LLM uncertainty into input-level, parameter-level, token-level, and decoding-process sources. Correspondingly, we categorize existing UQ methods into Bayesian, ensemble, consensus-based, and single-pass approaches. Furthermore, we introduce a comprehensive evaluation framework covering diverse generation settings and metrics. We empirically evaluate 21 typical UQ methods across three prominent LLM families, including Qwen3, Llama 3.2, and DeepSeek-V3, on benchmarks such as TriviaQA, GSM8K, and HumanEval. Our experimental results demonstrate that (i) the effectiveness of UQ methods is sensitive to task types and generation settings; (ii) consensus-based methods, typed Deg and EigV, consistently outperform other UQ approaches; and (iii) larger model scales correlate with lower uncertainty estimates, suggesting an empirical scaling law for LLM uncertainty. This work bridges the gap between theoretical origins and practical deployment, providing a versatile diagnostic tool for systematically quantifying uncertainty in LLM applications.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!