2605.28365v1 May 27, 2026 cs.AI

자연어 기반 수학적 추론에서 위험 관리 기반의 Lean-as-Judge 방법

Risk-Controlled Lean-as-Judge for Natural-Language Mathematical Reasoning

Haitham Bou-Ammar
Haitham Bou-Ammar
Citations: 2,607
h-index: 30
Rasul Tutunov
Rasul Tutunov
Citations: 781
h-index: 14
Matthieu Zimmer
Matthieu Zimmer
Citations: 70
h-index: 5
Xiaotong Ji
Xiaotong Ji
Citations: 39
h-index: 4
P. Bourigault
P. Bourigault
Citations: 23
h-index: 2

Lean은 자연어 형태의 수학적 답변을 평가하는 데 점점 더 많이 사용되고 있지만, 그 결과는 부분적인 정보만을 제공합니다. 많은 답변이 형식화되지 않으며, 증명 실패는 잘못된 유형의 문장이나 누락된 라이브러리 사실을 반영할 수 있으며, 반드시 틀린 답이라고 할 수 없습니다. MATH-500 데이터셋에서 우리는 (i) 이 신호가 증명 범위에 따라 크게 달라지며, 즉, 높은 증명 범위를 가질 때 정답인 경우가 96%이지만, 낮은 경우에는 20%에 불과하고, (ii) 이 신호는 희소하며 종종 정확하지 않다는 것을 보여줍니다. 7B 자동 형식화 도구는 문제의 28%에 대해서만 해를 찾았으며, 수동 검토 결과 해당 증명 중 약 43%만이 정확한 것으로 판명되었습니다. 우리는 Lean 추적 진단 정보를 기반으로 선택적인 위험 제어 경계를 갖는 COVCAL이라는 방법을 제안합니다. 이 방법은 허용된 답변에 대해 유한한 샘플 크기의 선택적 위험 경계를 적용하거나, 그렇지 않은 경우 결과를 제시하지 않습니다 (두 가지 방식: 보수적인 Bonferroni 경계 및 더 엄격한 개발-검증 규칙). 실행 가능성은 자동 형식화 범위에 따라 달라집니다. 7B 형식화 도구를 사용할 경우 신호가 너무 희소하여 모든 20개의 부트스트랩 파티션에서 Bonferroni 방식으로 결과를 제시하지 못합니다. 반면, 특정 증명기에 특화된 형식화 도구는 79%의 범위를 달성하고 이를 통해 20개 중 17개에서 실행 가능하게 만들어 약 48%의 문제를 98%의 정확도로 해결합니다. 자체 일관성만으로도 이미 91%의 정확도를 보이므로, 우리의 기여는 부분적인 형식화 신호가 언제, 그리고 어떤 형식화 도구를 사용할 때 위험 제어를 통해 신뢰할 수 있는지를 명확하게 설명하는 것입니다.

Original Abstract

Lean is increasingly used to judge natural-language mathematical answers, but its signal is partial: many answers never formalize, and a failed proof may reflect an ill-typed statement or a missing library fact, not a wrong answer. On MATH-500 we show this signal is (i) sharply coverage-dependent, that is the proof-winning answer is correct 96% of the time at high proved coverage but 20% at low, and (ii) sparse and often unfaithful: a 7B autoformalizer proves a class for only 28% of problems, and a manual audit finds only approximately 43% of those proofs faithful. We propose COVCAL, a selector over Lean-trace diagnostics that certifies a finite-sample selective-risk bound on accepted answers or abstains, under two regimes (a conservative Bonferroni bound and a tighter dev-then-cal rule). Feasibility depends on autoformalization coverage: with the 7B formalizer the signal is too sparse and Bonferroni abstains on all 20 bootstrap partitions, whereas a prover-specialized formalizer reaches 79% coverage and flips it to feasible on 17 of 20, accepting approximately 48% of problems at 0.98 accepted accuracy. Since self-consistency alone is already 91% accurate, our contribution is a precise account of when, and with which formalizer, a partial formal signal can be trusted under risk control.

0 Citations
0 Influential
15 Altmetric
75.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!