CALIBER: 언어 모델에서 추론 전후의 신뢰도 보정
CALIBER: Calibrating Confidence Before and After Reasoning in Language Models
추론 능력을 갖춘 언어 모델은 점점 더 어려운 질문에 답하는 것뿐만 아니라, 성공 가능성을 예측하도록 요구받고 있습니다. 기존 방법들은 일반적으로 신뢰도를 한 번만 측정합니다: 생각하기 전 또는 답변 후에. 우리는 추론 모델의 신뢰도가 상태 의존적이라고 주장합니다. 즉, 생각하기 전에는 모델이 프롬프트를 정확하게 해결할 확률을 추정해야 하고, 생각한 후에는 실제로 생성된 답변이 맞을 가능성을 예측해야 합니다. 이러한 구별은 적절한 지도 목표를 결정합니다. 즉, 프롬프트를 본 후에 생성된 신뢰도 추정은 프롬프트 수준의 성공 여부에 의해 지도되고, 답변 후에 생성된 신뢰도 추정은 개별 답변의 정확성에 의해 지도됩니다. 우리는 CALIBER(Calibration Before and After Reasoning)을 소개하며, 이는 두 가지 신뢰도 추정을 모두 수행하고, 각 추정에 해당 정보 상태에 맞는 목표를 사용하여 지도를 합니다. 이러한 통합 프로토콜 하에서, CALIBER는 7B 모델의 경우 BigMathDigits 데이터셋에서 가장 강력한 단일 신뢰도 기준보다 Expected Calibration Error (ECE)를 52.5% 줄입니다. 또한, CALIBER는 최고의 Brier score와 AUROC 값을 달성하며, 정확도 측면에서도 최상위 수준에 근접합니다. 더 큰 30B 모델에서, CALIBER는 BigMathDigits 데이터셋에서 가장 낮은 ECE를 달성하는 동시에 Brier score 및 AUROC에서 경쟁력 있는 성능을 보입니다. Out-of-distribution 환경에서는 GPQA 및 TriviaQA 데이터셋에서 최고의 ECE와 Brier score를 달성하며, SimpleQA에서도 경쟁력 있는 성능을 유지합니다. 추가적인 분석 결과, 이러한 위치-목표 정렬은 분포 변화가 발생하는 경우에 가장 큰 효과를 발휘하며, 모든 out-of-distribution 벤치마크에서 보정 오류를 지속적으로 줄입니다.
Reasoning language models are increasingly asked not only to answer difficult questions, but also to estimate their likelihood of success. Existing methods typically elicit confidence only once: either before thinking or after answering. We argue that confidence in reasoning models is state-dependent: before thinking, confidence should estimate the chance of the model correctly solving the prompt, while after thinking it should predict whether the realized answer is likely to be correct. This distinction determines the appropriate supervision target: prompt-level success should supervise confidence estimates made after seeing the prompt, while individual answer-level correctness should supervise confidence estimates made after answering. We introduce CALIBER (Calibration Before and After Reasoning), which elicits both estimates and supervises each with the target matched to its information state. Under this unified protocol, CALIBER reduces Expected Calibration Error (ECE) by 52.5% over the strongest single-confidence baseline on BigMathDigits for the 7B model, while achieving the best Brier score and AUROC, and remains within 2.1 points of the best accuracy. Further, on a larger 30B model, CALIBER achieves the best ECE on BigMathDigits while remaining competitive in Brier score and AUROC. Out of distribution, it achieves the best ECE and Brier score on GPQA and TriviaQA, and remains competitive on SimpleQA. Ablations further show that this position-target alignment is most beneficial under distribution shift where it consistently reduces calibration error across all out-of-distribution benchmarks.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.