2602.13540v1 Feb 14, 2026 cs.CL

대규모 언어 모델의 교정 연구: 응답에서 능력으로

On Calibration of Large Language Models: From Response To Capability

Sin-Han Yang
Sin-Han Yang
Citations: 17
h-index: 2
Cheng-Kuang Wu
Cheng-Kuang Wu
Citations: 301
h-index: 9
Chengxi Wu
Chengxi Wu
Citations: 88
h-index: 2
Chieh-Yen Lin
Chieh-Yen Lin
Citations: 176
h-index: 6
Yun-Nung Chen
Yun-Nung Chen
Citations: 223
h-index: 7
Hung-yi Lee
Hung-yi Lee
Citations: 213
h-index: 7
Shao-Hua Sun
Shao-Hua Sun
Citations: 122
h-index: 4

대규모 언어 모델(LLM)은 다양한 문제 해결 도구로 널리 사용되며, 신뢰성 있는 사용을 위해서는 정확한 신뢰도 추정이 매우 중요합니다. 기존의 LLM 교정 연구는 주로 응답 수준의 신뢰도에 초점을 맞추는데, 이는 생성된 단일 응답의 정확성을 추정하는 방식입니다. 그러나 이러한 방식은 모델이 전체적으로 얼마나 문제를 해결할 가능성이 있는지에 대한 핵심 질문을 다루지 못하며, 실제 환경과 괴리가 있습니다. 본 연구에서는 이러한 불일치가 현대 LLM 디코딩의 확률적 특성으로 인해 발생하는 현상임을 보여줍니다. 즉, 단일 응답의 정확성은 모델의 근본적인 능력 수준을 제대로 반영하지 못합니다. 이러한 문제를 해결하기 위해, 우리는 모델이 특정 질문에 대해 얼마나 정확하게 답변할 수 있는지를 목표로 하는 '능력 교정'을 제안합니다. 우리는 '능력 교정'과 '응답 교정'을 이론적, 실증적으로 명확히 구분하고, 두 가지가 서로 다르다는 것을 보여줍니다. 우리는 실증적인 평가 환경을 구축하고 다양한 신뢰도 추정 방법을 연구했습니다. 연구 결과는 '능력 교정'을 통해 신뢰도 예측 정확도(pass@$k$)와 추론 예산 할당을 개선할 수 있음을 보여주며, 이는 다양한 응용 분야에 잠재력을 가진 중요한 기반을 제공합니다.

Original Abstract

Large language models (LLMs) are widely deployed as general-purpose problem solvers, making accurate confidence estimation critical for reliable use. Prior work on LLM calibration largely focuses on response-level confidence, which estimates the correctness of a single generated output. However, this formulation is misaligned with many practical settings where the central question is how likely a model is to solve a query overall. We show that this mismatch results from the stochastic nature of modern LLM decoding, under which single-response correctness fails to reflect underlying model capability. To address this issue, we introduce capability calibration, which targets the model's expected accuracy on a query. We formally distinguish capability calibration from response calibration and show that the two differ both theoretically and empirically. We establish an empirical evaluation setup and study a range of confidence estimation methods. Our results demonstrate that capability-calibrated confidence improves pass@$k$ prediction and inference budget allocation, establishing a foundation with potential for diverse applications.

3 Citations
0 Influential
4.5 Altmetric
25.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!