2606.18021v1 Jun 16, 2026 cs.AI

LegalHalluLens: 유형별 환각 감사 및 신뢰성 있는 법률 AI를 위한 교정된 다중 에이전트 토론

LegalHalluLens: Typed Hallucination Auditing and Calibrated Multi-Agent Debate for Trustworthy Legal AI

L. Yadav
L. Yadav
Citations: 17
h-index: 2
Akshaj Gurugubelli
Akshaj Gurugubelli
Citations: 0
h-index: 0

법률 업무에 사용되는 AI 시스템은 평균적으로 약 52%의 오류율을 보이지만, 이러한 평균 수치는 오류가 집중되는 영역과 방향을 가리고 있어, 규정 준수 담당자가 신뢰성 있는 배포를 위한 실질적인 정보를 얻기 어렵습니다. 본 논문에서는 LegalHalluLens라는 감사 프레임워크를 제시합니다. 이 프레임워크는 세 가지 구성 요소로 이루어져 있습니다. 첫째, CUAD 데이터셋(Hendrycks et al., 2021)을 기반으로 숫자, 시간, 의무/권리, 사실 정보의 네 가지 법률 관련 주장에 대한 유형별 환각 프로필입니다. 둘째, 누락과 날조 편향을 단일 배포 비교 가능한 지표인 위험 방향 지수(RDI)로 줄이는 방법입니다. 셋째, 크기 및 방향 모두를 고려하여 교정된 유형별 토론 파이프라인입니다. 510개의 계약서와 249,252개의 조항 수준 데이터를 분석한 결과, 모델 내부적으로 의무/숫자 정보와 시간 정보 간에 약 38-40%p의 격차가 존재하며, 이러한 격차는 전체 보고에서 숨겨진다는 것을 확인했습니다. 또한, 동일한 52% 오류율을 가진 두 시스템이라도 서로 다른 RDI 값을 가질 수 있습니다. 토론 파이프라인은 허위 탐지를 45% 줄이며, 범주별 성능 향상은 진단 결과와 일치합니다. 또한, 상용 API와 비교하여 훨씬 작은 규모(40억 개의 활성 매개변수)의 백본을 사용합니다. 유형별 프로필과 RDI는 전체 지표에서 숨겨진 오류 패턴을 드러냅니다. 이러한 진단 정보는 다중 에이전트 토론 파이프라인의 교정 입력으로 사용될 수 있으며, 측정된 오류 패턴을 대상으로 하는 비대칭 게이트를 가진 Skeptic 시스템은 일반적인 튜닝을 거친 토론 시스템보다 우수한 성능을 보입니다. 본 프레임워크는 실제 법률 AI 배포에 대한 방향성을 고려한 조달, 책임 추적 및 에이전트 설계 기능을 지원합니다.

Original Abstract

AI systems deployed in legal workflows hallucinate at rates that aggregate metrics report at ~52%, but this average conceals where errors concentrate and in which direction they run, leaving compliance officers without an actionable signal for trustworthy deployment. We present LegalHalluLens, an auditing framework with three components: typed hallucination profiles across four legally-motivated claim categories (numeric, temporal, obligation/entitlement, factual) over CUAD (Hendrycks et al., 2021); a Risk Direction Index (RDI) that reduces omission-versus-invention bias to a single deployment-comparable scalar; and a typed debate pipeline calibrated to both magnitudes and directions. Across 510 contracts and 249,252 clause-level instances we measure a within-model gap of approximately 38-40 pp between obligation/numeric and temporal claims that aggregate reporting hides, and show that two systems with matched 52% rates can carry opposite RDIs. The debate pipeline reduces fabricated detections by 45% with per-category gains tracking the diagnosis, matching commercial APIs with a substantially smaller backbone (4B active parameters). Typed profiles and RDI surface failure modes that aggregate metrics hide; we further show these diagnostics serve as calibration inputs for multi-agent debate pipelines, where Skeptic challenges and asymmetric gates targeted at measured failure modes outperform generically-tuned debate. The framework supports direction-aware procurement, accountability, and agent design for legal AI deployed in the wild.

0 Citations
0 Influential
1 Altmetric
5.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!