추론을 통한 불확실성 추정: LLM의 다국어 및 교차 언어 MCQA 성능에 대한 대규모 연구
Estimating Uncertainty from Reasoning: A Large-Scale Study of Multi- and Crosslingual MCQA Performance in LLMs
불확실성 추정(UE)은 LLM 기반 시스템이 답변을 회피해야 할 때를 인식하도록 하지만, 기존 연구는 주로 영어에 초점을 맞추었습니다. 본 논문에서는 고, 중, 저 자원 환경을 포괄하는 22개 언어에 걸쳐 UE 방법론에 대한 최초의 대규모 평가를 제시합니다. 인간이 선별한 두 가지 질의응답 데이터 세트를 사용하여 다양한 모델 크기와 아키텍처에서 개방형 및 폐쇄형 UE 방법(총 9가지)을 비교하고, LLM을 심판으로 사용하거나 임베딩 기반 점수를 사용하는 방식을 피하여 평가 과정에서의 노이즈를 줄였습니다. 본 연구에서는 세 가지 주요 실질적인 결과를 보고합니다. 첫째, 저 자원 언어로 질문을 제시하면서 모델에게 영어로 추론하도록 유도하면 UE 성능이 크게 향상되는 것을 발견했는데, 이는 저 자원 언어에 대한 이해가 대체로 유지되고 있으며, 신뢰성 문제가 생성 단계에서 발생하는 것임을 시사합니다. 둘째, 모델에게 영어로 추론하도록 유도하면 저 자원 및 고 자원 언어 간의 UE 성능 격차가 줄어드는 것을 확인했는데, 이는 질문 언어보다 생성 언어가 더 중요하다는 것을 보여줍니다. 셋째, UE 방법 선택은 모델 크기에 따라 달라져야 합니다. 작은 규모의 모델에서는 개방형 확률 기반 방법이 다른 방법에 비해 우수하지만, 큰 규모의 모델에서는 폐쇄형 자기 설명 불확실성이 더 뛰어난 성능을 보입니다. 마지막으로, 선택적 예측을 위한 임계값 선택에 대한 분석을 제공하여 다국어 환경에서 답변 회피를 조정하는 데 필요한 지침을 제시합니다.
Uncertainty estimation (UE) enables LLM-powered systems to recognize when to abstain, yet existing research has predominantly focused on English. We present the first large-scale evaluation of UE methods across 22 languages, spanning high-, mid-, and low-resource settings. Using two human-curated Q\&A datasets, we compare open and closed box UE methods (nine in total) across different model sizes and architectures while eliciting long-form reasoning, avoiding LLM-as-a-judge and embedding-based scoring, which can introduce evaluation noise. We report three main actionable findings. First, we find that prompting models to reason in English while keeping questions in low-resource languages substantially improves UE performance, suggesting that comprehension of low-resource languages is largely intact, and that the reliability bottleneck lies in generation rather than understanding. Second, prompting models to reason in English closes the UE performance gap between low and high-resource languages, demonstrating that generation language matters more than the question language. Third, the choice of UE method should depend on model scale: at smaller scales, open-box probability-based methods outperform alternatives; at larger scales, closed-box self-verbalized uncertainty becomes superior. Finally, we provide an analysis of threshold selection for selective prediction, offering guidance on calibrating abstention in multilingual settings.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.