2602.07253v1 Feb 06, 2026 cs.AI

분포 외 탐지에서 환각 탐지로: 기하학적 관점

From Out-of-Distribution Detection to Hallucination Detection: A Geometric View

Litian Liu
Litian Liu
Citations: 164
h-index: 5
Reza Pourreza
Reza Pourreza
Citations: 61
h-index: 4
Yubin Jian
Yubin Jian
Citations: 55
h-index: 4
Roland Memisevic
Roland Memisevic
Citations: 94
h-index: 5
Ya-Qin Zhang
Ya-Qin Zhang
Citations: 86
h-index: 3

대규모 언어 모델에서 발생하는 환각(hallucination)을 탐지하는 것은 안전성과 신뢰성에 중요한 영향을 미치는 중요한 연구 과제입니다. 기존의 환각 탐지 방법은 질의 응답 작업에서 뛰어난 성능을 보이지만, 추론 능력을 요구하는 작업에서는 효과가 떨어지는 경우가 많습니다. 본 연구에서는 환각 탐지를 컴퓨터 비전 분야에서 널리 연구되는 분포 외(out-of-distribution, OOD) 탐지의 관점에서 재검토합니다. 언어 모델의 다음 토큰 예측을 분류 문제로 간주함으로써, OOD 기법을 적용할 수 있으며, 이를 위해 대규모 언어 모델의 구조적 차이를 고려한 적절한 수정이 필요합니다. 본 연구에서는 OOD 기반 접근 방식이 훈련 없이 단일 샘플을 사용하여 추론 작업에서 환각 탐지에 높은 정확도를 달성함을 보여줍니다. 전반적으로, 본 연구는 환각 탐지를 OOD 탐지로 재구성하는 것이 언어 모델의 안전성을 향상시키는 유망하고 확장 가능한 방법을 제시한다는 것을 시사합니다.

Original Abstract

Detecting hallucinations in large language models is a critical open problem with significant implications for safety and reliability. While existing hallucination detection methods achieve strong performance in question-answering tasks, they remain less effective on tasks requiring reasoning. In this work, we revisit hallucination detection through the lens of out-of-distribution (OOD) detection, a well-studied problem in areas like computer vision. Treating next-token prediction in language models as a classification task allows us to apply OOD techniques, provided appropriate modifications are made to account for the structural differences in large language models. We show that OOD-based approaches yield training-free, single-sample-based detectors, achieving strong accuracy in hallucination detection for reasoning tasks. Overall, our work suggests that reframing hallucination detection as OOD detection provides a promising and scalable pathway toward language model safety.

1 Citations
0 Influential
2.5 Altmetric
13.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!