2606.12900v1 Jun 11, 2026 cs.AI

인간과 유사한 기준 탐색을 통한 제로소스 LLM 환각 현상 감지

Zero-source LLM Hallucination Detection with Human-like Criteria Probing

Mingkui Tan
Mingkui Tan
Citations: 85
h-index: 5
Jiahao Yang
Jiahao Yang
Citations: 80
h-index: 3
Shuhai Zhang
Shuhai Zhang
Citations: 296
h-index: 10
Hailong Kang
Hailong Kang
Citations: 14
h-index: 2
Feng Liu
Feng Liu
Citations: 1
h-index: 1
Qi Chen
Qi Chen
Citations: 246
h-index: 7

대규모 언어 모델(LLM)은 사실과 다르거나 일관성이 없는 내용을 생성하여 환각 현상을 일으키는 경우가 많으며, 이는 안전한 사용에 상당한 위험을 초래합니다. 특히 제로소스 환경에서는 모델 내부 정보나 외부 참조가 전혀 없기 때문에 텍스트 기반의 질문-응답 쌍만을 이용하여 환각 현상을 감지하는 것이 매우 어렵습니다. 본 논문에서는 인간 평가자의 다면적인 추론 방식을 모방하는 환각 현상 감지 패러다임인 Human-like Criteria Probing for Hallucination Detection (HCPD)을 제안합니다. HCPD의 핵심은 Human-like Criteria Probbing (HCP) 메커니즘으로, LLM 에이전트가 판단을 해석 가능한 기준들의 가중치 합으로 분해하고, 각 기준별 점수를 종합하여 최종적인 진실성 지표를 산출합니다. 이러한 적응적 기능을 구현하기 위해 의미 일관성으로부터 얻은 약한 형태의 지도만을 사용하여 보상 기반 정렬 체계를 도입했습니다. 추론 단계에서는 강력한 의사 결정을 내리면서도 완전한 해석 가능성을 유지하기 위해 다중 샘플링 집계 전략을 사용합니다. 또한, 본 연구의 접근 방식의 신뢰성을 뒷받침하는 이론적 분석을 제공합니다. 광범위한 실험 결과, HCPD는 최첨단 기준 모델보다 일관되게 우수한 성능을 보이며, 제로소스 환경에서의 환각 현상 감지를 위한 효과적이고 설명 가능한 솔루션을 제시합니다. 관련 코드는 https://github.com/TRISKEL10N/HCPD 에서 확인할 수 있습니다.

Original Abstract

Large language models (LLMs) often hallucinate by generating factually incorrect or unfaithful content, posing significant risks to their safe use. Detecting such hallucinations is particularly challenging under the zero-source constraint, where no model internals or external references are available, and detection must rely solely on the textual query-answer pair. In this paper, we propose Human-like Criteria Probing for Hallucination Detection (HCPD), a paradigm that emulates the multi-faceted reasoning of human evaluators. Its core is a Human-like Criteria Probing (HCP) mechanism, in which a LLM agent adaptively decomposes its judgment into a weighted set of interpretable criteria and aggregates criterion-specific scores into a final truthfulness measure. To achieve this adaptive capability, we introduce a reward-based alignment scheme using only weak supervision from semantic consistency. At inference, we employ a multi-sampling aggregation strategy to ensure robust decisions while preserving full interpretability. We further provide theoretical analysis supporting the reliability of our approach. Extensive experiments show that HCPD consistently outperforms state-of-the-art baselines, offering an effective and explainable solution for zero-source hallucination detection. Code is available at https://github.com/TRISKEL10N/HCPD.

0 Citations
0 Influential
30.493061443341 Altmetric
0.0 Score
Original PDF
2

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!