DeepFact: 심층 연구 사실성 검증을 위한 공동 진화 벤치마크 및 에이전트
DeepFact: Co-Evolving Benchmarks and Agents for Deep Research Factuality
검색 기반 LLM 에이전트는 심층 연구 보고서(DRR)를 생성할 수 있지만, 주장 수준의 사실성을 검증하는 것은 여전히 어려운 과제입니다. 기존의 사실 검증 도구는 주로 일반적인 영역의 사실 기반적인 주장을 대상으로 설계되었으며, 이러한 검증 도구가 DRR에 적용될 수 있는지 여부를 테스트할 수 있는 벤치마크는 존재하지 않습니다. 하지만 그러한 벤치마크를 구축하는 것 자체가 어렵습니다. 우리는 먼저 전문가가 레이블링한 정적 벤치마크가 이 환경에서 취약하다는 것을 보여줍니다. 통제된 연구에서 박사 학위 소지자 전문가들은 지원 없이 검증 가능한 주장들의 숨겨진 미세 수준 데이터셋에 대해 60.8%의 정확도를 달성하는 데 그쳤습니다. 우리는 '감사 후 점수(Audit-then-Score, AtS)'라는 공동 진화 벤치마킹 방식을 제안합니다. AtS에서는 벤치마크 레이블과 근거가 명시적으로 수정 가능합니다. 검증 도구가 현재 벤치마크와 의견이 다를 경우, 증거를 제시해야 하며, 감사자가 분쟁을 조정하고, 수용된 수정 사항은 모델이 평가되기 전에 벤치마크를 업데이트합니다. 네 번의 AtS 반복을 통해, 전문가의 미세 수준 정확도는 90.9%로 상승했습니다. 이는 전문가가 일회성 레이블러로서보다 감사자로서 훨씬 더 신뢰할 수 있음을 나타냅니다. 우리는 AtS를 'DeepFact-Bench'라는 버전 관리 DRR 사실성 벤치마크(감사 가능한 근거 포함)와 'DeepFact-Eval'이라는 문서 수준의 검증 에이전트(그룹화된 경량 버전 포함)로 구현했습니다. DeepFact-Eval은 DeepFact-Bench에서 기존 검증 도구보다 뛰어난 성능을 보이며, 외부 사실성 데이터셋에도 잘 적용됩니다.
Search-augmented LLM agents can produce deep research reports (DRRs), but verifying claim-level factuality remains challenging. Existing fact-checkers are primarily designed for general-domain, factoid-style atomic claims, and there is no benchmark to test whether such verifiers transfer to DRRs. Yet building such a benchmark is itself difficult. We first show that static expert-labeled benchmarks are brittle in this setting: in a controlled study with PhD-level specialists, unassisted experts achieve only 60.8% accuracy on a hidden micro-gold set of verifiable claims. We propose Evolving Benchmarking via Audit-then-Score (AtS), where benchmark labels and rationales are explicitly revisable: when a verifier disagrees with the current benchmark, it must submit evidence; an auditor adjudicates the dispute; and accepted revisions update the benchmark before models are scored. Across four AtS rounds, expert micro-gold accuracy rises to 90.9%, indicating experts are substantially more reliable as auditors than as one-shot labelers. We instantiate AtS as DeepFact-Bench, a versioned DRR factuality benchmark with auditable rationales, and DeepFact-Eval, a document-level verification agent (with a grouped lite variant) that outperforms existing verifiers on DeepFact-Bench and transfers well to external factuality datasets.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.