2608.14210v1 Aug 14, 2026 cs.CL

법률 분야의 검색 기반 생성(RAG) 시스템은 여전히 얼마나 많은 환각 현상을 보이는가?

How Much Do Legal RAG Systems Still Hallucinate?

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
S. Abualhaija
S. Abualhaija
Citations: 462
h-index: 11
D. Bianculli
D. Bianculli
Citations: 54
h-index: 5

환각 현상은 법률 분야에서 검색 기반 생성(RAG) 시스템의 주요 과제로, 근거 없는 답변은 심각한 결과를 초래할 수 있습니다. 이 문제를 더 잘 이해하기 위해, 우리는 GDPR(영어) 및 국내 민법(프랑스어)이라는 두 가지 법률 데이터셋에 대한 8개의 RAG 시스템에서 환각 현상의 세부적인 분석을 수행했습니다. 주장 수준과 답변 수준의 평가를 통해, 우리는 환각 발생 빈도와 심각도를 보고하고, 질문 유형 및 사용자 페르소나별 성능을 분석했으며, 법률 전문가가 작성한 독립적인 142개의 질문에 대한 결과를 검증했습니다. 우리의 결과는 환각 현상이 여전히 광범위하게 나타나는 것으로 보여주며, 가장 성능이 좋은 시스템에서는 응답의 10% 미만에서, 최악의 경우 거의 절반에서 환각 현상이 발생합니다. 또한, 잘못된 전제를 포함하는 질문(즉, 오류를 포함하고 거부해야 하는 가정)은 수동으로 작성된 질문에서 높은 환각 발생률을 보이는 것으로 나타났습니다.

Original Abstract

Hallucination is a major challenge for retrieval-augmented generation (RAG) systems in the legal domain, where ungrounded answers can lead to serious consequences. To better understand this problem, we conduct a fine-grained analysis of hallucination behavior in eight legal RAG systems across two legal corpora, the GDPR (in English) and a national civil law (in French). Using claim-level and answer-level evaluation, we report on hallucination density and severity, analyze performance across question categories and user personas, and validate our findings on an independent set of 142 legal-expert-authored questions. Our results show that hallucinations remain pervasive, ranging from less than 10% of responses for the best-performing systems to nearly half in the worst case. We further find that false-premise questions, containing incorrect assumptions that must be rejected, produce high hallucination rates on the manually-drafted questions.

0 Citations
0 Influential
5.5 Altmetric
27.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!