대규모 언어 모델에서의 신뢰도의 계산적 기반
The Computational Basis of Confidence in Large Language Models
언어 모델의 신뢰성 있는 활용을 위해서는 모델 스스로 제시하는 답변의 정확성에 대한 확률, 즉 '신뢰도'가 필수적입니다. 기존 연구는 주로 신뢰도가 얼마나 정확하게 정답 여부를 예측하고, 얼마나 잘 보정되어 있는지에 초점을 맞추었지만, '신뢰도 신호' 자체가 무엇을 의미하는지에 대한 근본적인 질문은 여전히 열려 있었습니다. 답변 로짓(logits)은 정상적인 신뢰도를 계산하기에 충분한 잠재적 의사 결정 변수를 반영할 수도 있고, 또는 사용 가능한 증거를 비-베이즈 방식으로 결합하는 휴리스틱한 선호도 신호를 나타낼 수도 있습니다. 본 연구에서는 계산 신경과학 분야의 규범적 프레임워크인 통계적 의사 결정 신뢰도(Statistical Decision Confidence, SDC)를 활용하여 이 문제를 해결하고자 합니다. 답변 로짓과 관련된 차이(LD)를 잠재적 의사 결정 변수의 지표로 간주하고, SDC에서 예측하는 특징들을 검증했습니다. 세 가지 인지 구분 작업 및 기억 기반 의사 결정 작업을 수행했으며, 여기에는 세 개의 다중 모드 비-추론 모델과 하나의 추론 모델이 포함되었습니다. 결과적으로 LD는 이러한 특징들, 특히 정확/오류에 대한 '접힌 X' 패턴을 만족했으며, 이는 해당 환경에서 답변 로짓이 휴리스틱한 선호도 점수가 아닌 잠재적 의사 결정 변수의 단조적인 지표로 작동한다는 것을 보여줍니다. 복잡한 시각적 추론에서는 LD가 객관적인 작업 난이도를 넘어 정답 여부를 예측하는 경향을 보였지만, SDC의 완전한 기하학적 특징은 나타나지 않았습니다. 이는 명시적인 규범적 과정 모델이 없을 때 프레임워크의 현재 한계를 보여줍니다. 본 연구는 다중 모드 언어 모델에서 신뢰도에 대한 계산적 설명을 제공하고, 답변 로짓이 잠재적 의사 결정 변수의 지표로 작용하는 조건을 구체적으로 제시하며, 생물학 및 인공지능 시스템 전반의 신뢰도 연구를 위한 통합 프레임워크로서 SDC를 확립합니다.
Reliable confidence -- the probability that a model's own answer is correct -- is essential for the trustworthy deployment of language models. Existing work has largely evaluated confidence by how well it predicts correctness and whether it is calibrated, leaving open a more fundamental question: what does the confidence signal itself represent? Answer logits may reflect a latent decision variable sufficient to compute normative confidence, or instead a heuristic preference signal that combines the available evidence in a non-Bayesian manner. We address this using statistical decision confidence (SDC), a normative framework from computational neuroscience. Treating the answer-logit difference (LD) as a candidate readout of the latent decision variable, we test the qualitative signatures predicted by SDC. Across three perceptual discrimination tasks and a memory-based decision task, spanning three multimodal non-reasoning models and one reasoning model, LD satisfied these signatures -- including the diagnostic correct/error folded-X pattern -- showing that, in these settings, answer logits behave as monotonic readouts of a latent decision variable rather than heuristic preference scores. In complex visual reasoning, LD continued to predict correctness beyond objective task difficulty, but the full geometric signatures of SDC were absent, illustrating the current boundary of the framework when explicit normative process models are unavailable. These results provide a computational account of confidence in multimodal language models, delineate when answer logits behave as readouts of a latent decision variable, and establish SDC as a unifying framework for studying confidence across biological and artificial intelligence.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.