2606.18557v1 Jun 17, 2026 cs.AI

DeFAb: 기초 모델에서의 반증 가능한 추론을 위한 검증 가능한 벤치마크

DeFAb: A Verifiable Benchmark for Defeasible Abduction in Foundation Models

Patrick Cooper
Patrick Cooper
Citations: 0
h-index: 0
Alvaro Velasquez
Alvaro Velasquez
Citations: 130
h-index: 5

본 연구에서 개발한 규칙 기반 논리 솔버는 우리의 벤치마크에 있는 모든 사례를 50 마이크로초 이내에 100% 정확도로 해결합니다. 반면, 현재 가장 성능이 뛰어난 언어 모델은 최대 65%의 정확도를 보이지만, 렌더링 강건성 평가에서는 23.5%까지 떨어집니다 (네 가지 표면 렌더링 방식을 고려한 최악의 경우). 우리는 DeFAb(Defeasible Abduction Benchmark)이라는 데이터셋과 생성 파이프라인을 소개합니다. 이는 지난 40년 동안 공개적으로 자금을 지원받아 구축된 지식 베이스를 형식적으로 기반으로 하는 사례로 변환하여, 기본 규칙을 무시하면서 관련 없는 기대를 유지하며 이상 현상을 설명하는 가설(반증 가능한 추론)을 구성하도록 합니다. 모든 가설은 유효한 도출, 보수성 및 최소성에 대한 다항 시간 검사를 통과해야 하므로, DeFAb는 논리적 엄밀성을 창의성과 이론적 추론을 측정하는 기준으로 사용하며, 유창하지만 이론을 파괴하는 텍스트가 아닌, 체계적인 이론 수정 과정을 평가합니다. 이 파이프라인은 분류 계층 구조(OpenCyc, YAGO, Wikidata)를 행동 속성 그래프(ConceptNet, UMLS)와 결합하여 18개 소스에서 추출한 3375만 개의 구체화된 규칙을 사용하여 372,648개 이상의 사례를 생성합니다. 이 데이터는 다항 시간으로 검증 가능한 표준 정답을 가진 세 가지 수준으로 구성됩니다. 현재 가장 성능이 뛰어난 네 모델은 반증 가능한 추론을 안정적으로 학습하지 못하는 것으로 나타났습니다. 렌더링 강건성 Level 2의 정확도는 7.8-23.5%이며, Chain-of-Thought 방식의 결과 편차는 약 36pp로, 모델 간 격차보다 큽니다. 또한, 매칭된 오염 제어 실험을 통해 Level 3에서 +19.4pp의 성능 차이가 있음을 확인했습니다. 우리는 더 높은 난이도의 DeFAb-Hard(235개의 Level 3 사례; 최적 모델은 53.3% vs 100% 기호 처리)와 CONJURE(커널 검증을 거친 혁신적인 창의성 변형으로, Lean 4/Mathlib 기반의 560개 사례로 구성되며, 표준 정답은 기존에 증명 커널이 포함하지 않던 정의이며, 심사관 없이 검증 가능합니다. 파일럿 테스트 결과 새로운 개념은 발견되지 않았습니다)를 추가적으로 공개합니다. 동일한 검증기는 강화 학습(DPO, RLVR/GRPO)을 위한 정확한 보상으로도 사용됩니다. DeFAb는 MIT 라이선스 하에 https://huggingface.co/datasets/PatrickAllenCooper/DeFAb 에서 제공됩니다.

Original Abstract

A rule-based logic solver resolves every instance in our benchmark in under 50 microseconds with 100% accuracy; the best frontier language model reaches 65% at best and drops to 23.5% under rendering-robust evaluation (worst case over four surface renderings). We introduce DeFAb (Defeasible Abduction Benchmark), a dataset and generation pipeline that converts four decades of publicly funded knowledge bases into formally grounded instances for defeasible abduction: constructing hypotheses that explain anomalies by overriding defaults while preserving unrelated expectations. Because every hypothesis must pass polynomial-time checks for valid derivation, conservativity, and minimality, DeFAb makes logical rigor the instrument for measuring creativity and theoretical reasoning, scoring the disciplined construction of theory revisions rather than fluent but theory-destroying prose. The pipeline pairs taxonomic hierarchies (OpenCyc, YAGO, Wikidata) with behavioral property graphs (ConceptNet, UMLS) to produce 372,648+ instances across 33.75M materialized rules from 18 sources, in three levels with polynomial-time verifiable gold standards. Four frontier models do not reliably internalize defeasible reasoning: rendering-robust Level 2 accuracy is 7.8-23.5%; chain-of-thought variance (~36 pp) exceeds any inter-model gap; and a matched contamination control isolates a +19.4 pp Level 3 gap. We further release DeFAb-Hard (a 235-instance Level 3 difficulty variant; best model 53.3% vs 100% symbolic) and CONJURE (a kernel-verified transformative-creativity variant of 560 Lean 4/Mathlib instances whose gold answers are definitions the proof kernel did not previously contain, judge-free verifier; a pilot finds zero novel concepts). The same verifier doubles as an exact reward for preference optimization (DPO, RLVR/GRPO). Released under MIT at https://huggingface.co/datasets/PatrickAllenCooper/DeFAb.

0 Citations
0 Influential
22.5 Altmetric
112.5 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!