증거 기반 검증된 자율 추론: 도구 인증을 통한 형식적 증명으로 LLM 환각 현상을 제거하는 방법
Evidence-Grounded Verified Agentic Reasoning: A Path Toward Eliminating LLM Hallucination in Empirical Inference via Tool-Attested Kernel Proofs
단순히 도구를 활용하는 것만으로는 LLM의 경험적 추론을 통제할 수 없습니다. 허용되는 출력은 반드시 검증된 증거에서 비롯되어야 하며, 받아들여지는 추론은 형식적인 검토를 견딜 수 있어야 합니다. 본 논문에서는 Lean 4 기반의 도구 호출 아키텍처인 EG-VAR (Evidence-Grounded Verified Agentic Reasoning)을 제안합니다. EG-VAR에서 Lean 커널은 도구 인증 공리와 명시된 소스 연결을 통해 검증된 주장을 생성하는 유일한 주체입니다. 모든 검증된 출력은 검증된 도구 호출(정리 3.1)과 커널에 의해 확인된 유효한 추론 체인(정리 3.2)에서 구조적으로 파생됩니다. 나머지 출력은 재현 가능한 감사 추적을 통해 '정보 없음'으로 처리됩니다. TableBench의 숫자 추론 데이터셋의 일부(n=120)를 사용하여 평가한 결과, EG-VAR는 120/120의 정확도를 달성했으며, 동일 도구를 사용하는 기준 모델은 95%의 정확도를 보였습니다. 반사실적 스트레스 테스트(5개 영역 x 2개의 모델)에서는 EG-VAR가 100%의 소스 충실도를 유지하는 반면, 동일 도구 모델은 80-90%로 감소하고, 도구를 사용하지 않는 모델은 50-80%에 그쳤습니다. LLM을 배포 시점의 형식화기로 사용할 때, Sonnet 모델에서 잔여 의미-형식화 오류는 3.3%, Opus 모델에서는 1.7%였습니다. EG-VAR는 고위험 분야의 경험적 주장을 위한 기술 거버넌스 인터페이스로 기능합니다. 공식적인 보조 시스템은 대상 제안, 소스 범위, 증거 경계, 증명 의무 및 회피 조건을 감사할 수 있도록 하여 현재 지원되지 않는 검증된 출력을 제거하고, 형식화 오류, 연결 및 소스 권한 분쟁, 모호성 및 회피를 명시적인 감사 대상으로 전환합니다. 시간이 지남에 따라 데이터셋, API, 공개 기록 및 AI 생성 문서에서 유형화된 보조 시스템을 활용하여 이러한 형식화 부담을 재사용 가능한 인프라로 전환할 수 있습니다.
Tool access alone does not make LLM empirical reasoning governable: accepted outputs need not descend from attested evidence, and accepted deductions need not hold up under formal scrutiny. We present EG-VAR (Evidence-Grounded Verified Agentic Reasoning), a Lean 4-based tool-calling architecture in which the Lean kernel is the sole minter of Verified claims via tool-attestation axioms and declared source lifts. Every verified output structurally descends from an attested tool call (Thm. 3.1) and a kernel-checked chain of valid inference (Thm. 3.2); residual outputs are honest Abstain with a replayable audit trail. On a subcollection of TableBench numerical reasoning (n=120), EG-VAR attains 120/120 versus a 95% same-tool baseline; on counterfactual stress tests (5 domains x 2 models), EG-VAR stays 100% source-faithful while same-tool drops to 80-90% (no-tool 50-80%). With the LLM as deployment-time formalizer, residual semantic-formalization error is 3.3% on Sonnet and 1.7% on Opus. We position EG-VAR as a technical-governance interface for high-stakes empirical claims: a formal sidecar makes the target proposition, source scope, evidence boundary, proof obligation, and abstention condition auditable, eliminating unsupported Verified outputs today while turning formalization errors, lift and source-authority disputes, ambiguities, and abstentions into explicit audit targets. Over time, typed sidecars in datasets, APIs, public records, and AI-generated documents can amortize this formalization burden into reusable infrastructure.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.