Draft-and-Prune: 논리적 추론을 위한 자동 형식화의 신뢰성 향상
Draft-and-Prune: Improving the Reliability of Auto-formalization for Logical Reasoning
자동 형식화(AF)는 자연어 추론 문제를 솔버가 실행할 수 있는 프로그램으로 변환하여, 심볼릭 솔버가 정확한 논리적 추론을 수행할 수 있도록 합니다. 하지만 현재 AF 파이프라인은 불안정하며, 프로그램이 실행에 실패하거나, 실행되더라도 부정확한 의미를 포함할 수 있습니다. 기존 연구에서는 주로 솔버 피드백을 기반으로 한 수정을 통해 구문 오류를 완화했지만, 의미 오류를 줄이는 것은 여전히 중요한 과제입니다. 본 논문에서는 다양성과 검증을 통해 AF 기반 논리적 추론을 향상시키는 추론 시간 프레임워크인 Draft-and-Prune (D&P)를 제안합니다. D&P는 먼저 여러 개의 자연어 계획을 생성하고, 프로그램 생성을 이러한 계획에 기반하여 수행합니다. 또한, 실행 가능하지만 모순적이거나 모호한 형식화를 제거하고, 살아남은 경로에서 예측값을 집계하여 다수결 방식으로 결과를 결정합니다. AR-LSAT, ProofWriter, PrOntoQA, LogicalDeduction의 네 가지 대표적인 벤치마크에서, D&P는 추가적인 감독 없이 AF 기반 추론의 성능을 크게 향상시켰습니다. AR-LSAT에서, AF만 사용하는 환경에서 D&P는 GPT-4를 사용하여 78.43%의 정확도를, GPT-4o를 사용하여 78.00%의 정확도를 달성했으며, 이는 가장 강력한 AF 기반 모델인 MAD-LOGIC과 CLOVER를 크게 능가하는 결과입니다. D&P는 다른 벤치마크에서도 거의 최고 수준의 성능을 달성했으며, PrOntoQA와 LogicalDeduction에서는 100%의 정확도를 보였습니다.
Auto-formalization (AF) translates natural-language reasoning problems into solver-executable programs, enabling symbolic solvers to perform sound logical deduction. In practice, however, AF pipelines are currently brittle: programs may fail to execute, or execute but encode incorrect semantics. While prior work largely mitigates syntactic failures via repairs based on solver feedback, reducing semantics failures remains a major bottleneck. We propose Draft-and-Prune (D&P), an inference-time framework that improves AF-based logical reasoning via diversity and verification. D&P first drafts multiple natural-language plans and conditions program generation on them. It further prunes executable but contradictory or ambiguous formalizations, and aggregates predictions from surviving paths via majority voting. Across four representative benchmarks (AR-LSAT, ProofWriter, PrOntoQA, LogicalDeduction), D&P substantially strengthens AF-based reasoning without extra supervision. On AR-LSAT, in the AF-only setting, D&P achieves 78.43% accuracy with GPT-4 and 78.00% accuracy with GPT-4o, significantly outperforming the strongest AF baselines MAD-LOGIC and CLOVER. D&P then attains near-ceiling performance on the other benchmarks, including 100% on PrOntoQA and LogicalDeduction.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.