2512.11614v3 Dec 12, 2025 cs.CL

환각 현상 제어: 언어 모델의 상호 정보량 경계를 위한 멀린-아서 프로토콜

Bounding Hallucinations: Merlin-Arthur Protocols for Mutual-Information Bounds in Language Models

Kristian Kersting
Kristian Kersting
Citations: 95
h-index: 5
Bjorn Deiseroth
Bjorn Deiseroth
Citations: 860
h-index: 9
Letitia Parcalabescu
Letitia Parcalabescu
Citations: 1
h-index: 1
Max Henning Höth
Max Henning Höth
Citations: 1
h-index: 1

검색 증강 생성(RAG)은 검색된 문맥 정보를 활용하여 대규모 언어 모델(LLM)을 안내하지만, 검색 결과를 검증 가능한 증거로 취급하지 않고 휴리스틱으로 간주합니다. 이는 근거 없는 답변, 환각 현상, 그리고 잘못된 문맥 정보에 대한 의존성을 야기할 수 있습니다. 본 연구에서는 RAG 파이프라인을 상호 작용적 증명 시스템으로 처리하기 위해 멀린-아서(M/A) 프로토콜을 적용하는 새로운 평가, 데이터 증강 및 훈련 프레임워크를 제안합니다. 아서(생성 LLM)는 출처가 불분명한 문맥 정보를 받고, 멀린은 유용한 증거를 제공하며, 모가나는 악의적이고 오해를 불러일으킬 수 있는 문맥 정보를 주입합니다. 우리는 이 프레임워크를 XAI 방법과 함께 구현하여 아서에게 가장 큰 영향을 미치는 증거를 스스로 평가하고 수정하도록 합니다. 본 연구에서는 설명 정확도(explanation fidelity)와 모델 예측 오류, 그리고 불완전한 벤치마크를 분리하는 '설명 정보 비율(Explained Information Fraction, EIF)' 점수를 제안하며, M/A 상호 정보량 하한을 현실적인 경험적 환경에 맞게 정규화합니다. 이러한 문맥 정보를 사용하여 학습된 아서는 증거가 답변을 뒷받침할 때 답변하고, 그렇지 않은 경우에는 답변하지 않도록 학습됩니다. 5개의 질의응답 벤치마크와 4가지 LLM(10억 개에서 320억 개 파라미터)에 대해, M/A는 충분한 문맥이 없는 경우 잘못된 답변을 최대 35%p 감소시키며, 일반적인 미세 조정보다 18-20%p 더 높은 성능을 보입니다. 수동으로 주석 처리된 불가능한 예제나 선호도 쌍이 없어도 '거절(abstain)' 기능이 나타납니다. M/A 학습 시 EIF 점수가 0.1-0.4 향상되며, 일반적인 미세 조정보다 0.33-0.38 더 높은 성능을 보입니다. 멀린/모가나 문맥 정보를 자동 하드 포지티브 및 네거티브 샘플로 재사용하여 검색기의 Recall@1 성능을 2%p 향상시켰습니다. 높은 정확도가 문맥에서 답변으로의 엔트로피 흐름을 보장하지는 않지만, 본 연구에서 제안하는 EIF 점수는 RAG 생성 모델에 대한 최초의 문맥-답변 간 상호 정보량 경계이며, 자율적인 상호 작용적 증명 스타일의 감독을 통해 검색된 문서가 검증 가능한 증거로 취급되는 RAG 시스템을 가능하게 합니다.

Original Abstract

Retrieval-augmented generation (RAG) relies on retrieved context to guide large language models (LLM), yet treats the retrieval as a heuristic rather than verifiable evidence -- leading to unsupported answers, hallucinations, and reliance on spurious context. We introduce a novel evaluation, data augmentation, and training framework that treats the RAG pipeline as an interactive proof system by adapting the Merlin-Arthur (M/A) protocol: Arthur (the generator LLM) receives context of unknown provenance and Merlin gives helpful evidence, while Morgana injects adversarial, misleading context. We implement them both with an XAI method to self-assess and modify evidence most influential to Arthur. Based on this framework we propose the Explained Information Fraction (EIF) score, that disentangles explanation fidelity from model predictive errors and imperfect benchmarks, and normalizes M/A mutual-information lower bounds to realistic empirical settings. When trained with those contexts, Arthur learns to answer when evidence supports the answer and abstain when evidence is insufficient. Across five QA benchmarks and four LLMs (1B to 32B parameters), M/A reduces incorrect answers under insufficient context by up to 35pp during training and by 18-20pp over vanilla finetuning. Abstention emerges even \emph{without any manually annotated unanswerable example or preference pair}. We improve EIF-cond by 0.1-0.4 during M/A training and by 0.33-0.38 over vanilla finetuning. Reusing Merlin/Morgana contexts as automatic hard positives and negatives also raises retriever Recall@1 by 2pp. While high accuracy does not guarantee entropy flow from context to answer, our EIF scores -- to our knowledge, the first context-to-answer information bound for RAG generators -- show that autonomous interactive-proof-style supervision enables RAG systems that treat retrieved documents as verifiable evidence.

0 Citations
0 Influential
4.5 Altmetric
22.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!