2603.17233v1 Mar 18, 2026 cs.AI

Draft-and-Prune: 논리적 추론을 위한 자동 형식화의 신뢰성 향상

Draft-and-Prune: Improving the Reliability of Auto-formalization for Logical Reasoning

Zhiyu Ni
Zhiyu Ni
Citations: 16
h-index: 2
Zhengwen Liang
Zhengwen Liang
Citations: 316
h-index: 2
Liang Song
Liang Song
Citations: 46
h-index: 4
Chenrui Cao
Chenrui Cao
Computer Science and Technology, University of Science and Technology of China
Citations: 34
h-index: 2
Xian Zhang
Xian Zhang
Citations: 178
h-index: 5
Alberto Sangiovanni-Vincentelli
Alberto Sangiovanni-Vincentelli
Citations: 47
h-index: 1
Pierluigi Nuzzo
Pierluigi Nuzzo
Citations: 7
h-index: 2

자동 형식화(AF)는 자연어 추론 문제를 솔버가 실행할 수 있는 프로그램으로 변환하여, 심볼릭 솔버가 정확한 논리적 추론을 수행할 수 있도록 합니다. 하지만 현재 AF 파이프라인은 불안정하며, 프로그램이 실행에 실패하거나, 실행되더라도 부정확한 의미를 포함할 수 있습니다. 기존 연구에서는 주로 솔버 피드백을 기반으로 한 수정을 통해 구문 오류를 완화했지만, 의미 오류를 줄이는 것은 여전히 중요한 과제입니다. 본 논문에서는 다양성과 검증을 통해 AF 기반 논리적 추론을 향상시키는 추론 시간 프레임워크인 Draft-and-Prune (D&P)를 제안합니다. D&P는 먼저 여러 개의 자연어 계획을 생성하고, 프로그램 생성을 이러한 계획에 기반하여 수행합니다. 또한, 실행 가능하지만 모순적이거나 모호한 형식화를 제거하고, 살아남은 경로에서 예측값을 집계하여 다수결 방식으로 결과를 결정합니다. AR-LSAT, ProofWriter, PrOntoQA, LogicalDeduction의 네 가지 대표적인 벤치마크에서, D&P는 추가적인 감독 없이 AF 기반 추론의 성능을 크게 향상시켰습니다. AR-LSAT에서, AF만 사용하는 환경에서 D&P는 GPT-4를 사용하여 78.43%의 정확도를, GPT-4o를 사용하여 78.00%의 정확도를 달성했으며, 이는 가장 강력한 AF 기반 모델인 MAD-LOGIC과 CLOVER를 크게 능가하는 결과입니다. D&P는 다른 벤치마크에서도 거의 최고 수준의 성능을 달성했으며, PrOntoQA와 LogicalDeduction에서는 100%의 정확도를 보였습니다.

Original Abstract

Auto-formalization (AF) translates natural-language reasoning problems into solver-executable programs, enabling symbolic solvers to perform sound logical deduction. In practice, however, AF pipelines are currently brittle: programs may fail to execute, or execute but encode incorrect semantics. While prior work largely mitigates syntactic failures via repairs based on solver feedback, reducing semantics failures remains a major bottleneck. We propose Draft-and-Prune (D&P), an inference-time framework that improves AF-based logical reasoning via diversity and verification. D&P first drafts multiple natural-language plans and conditions program generation on them. It further prunes executable but contradictory or ambiguous formalizations, and aggregates predictions from surviving paths via majority voting. Across four representative benchmarks (AR-LSAT, ProofWriter, PrOntoQA, LogicalDeduction), D&P substantially strengthens AF-based reasoning without extra supervision. On AR-LSAT, in the AF-only setting, D&P achieves 78.43% accuracy with GPT-4 and 78.00% accuracy with GPT-4o, significantly outperforming the strongest AF baselines MAD-LOGIC and CLOVER. D&P then attains near-ceiling performance on the other benchmarks, including 100% on PrOntoQA and LogicalDeduction.

1 Citations
0 Influential
2.5 Altmetric
13.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!