자연어 기반 수학적 추론에서 위험 관리 기반의 Lean-as-Judge 방법
Risk-Controlled Lean-as-Judge for Natural-Language Mathematical Reasoning
Lean은 자연어 형태의 수학적 답변을 평가하는 데 점점 더 많이 사용되고 있지만, 그 결과는 부분적인 정보만을 제공합니다. 많은 답변이 형식화되지 않으며, 증명 실패는 잘못된 유형의 문장이나 누락된 라이브러리 사실을 반영할 수 있으며, 반드시 틀린 답이라고 할 수 없습니다. MATH-500 데이터셋에서 우리는 (i) 이 신호가 증명 범위에 따라 크게 달라지며, 즉, 높은 증명 범위를 가질 때 정답인 경우가 96%이지만, 낮은 경우에는 20%에 불과하고, (ii) 이 신호는 희소하며 종종 정확하지 않다는 것을 보여줍니다. 7B 자동 형식화 도구는 문제의 28%에 대해서만 해를 찾았으며, 수동 검토 결과 해당 증명 중 약 43%만이 정확한 것으로 판명되었습니다. 우리는 Lean 추적 진단 정보를 기반으로 선택적인 위험 제어 경계를 갖는 COVCAL이라는 방법을 제안합니다. 이 방법은 허용된 답변에 대해 유한한 샘플 크기의 선택적 위험 경계를 적용하거나, 그렇지 않은 경우 결과를 제시하지 않습니다 (두 가지 방식: 보수적인 Bonferroni 경계 및 더 엄격한 개발-검증 규칙). 실행 가능성은 자동 형식화 범위에 따라 달라집니다. 7B 형식화 도구를 사용할 경우 신호가 너무 희소하여 모든 20개의 부트스트랩 파티션에서 Bonferroni 방식으로 결과를 제시하지 못합니다. 반면, 특정 증명기에 특화된 형식화 도구는 79%의 범위를 달성하고 이를 통해 20개 중 17개에서 실행 가능하게 만들어 약 48%의 문제를 98%의 정확도로 해결합니다. 자체 일관성만으로도 이미 91%의 정확도를 보이므로, 우리의 기여는 부분적인 형식화 신호가 언제, 그리고 어떤 형식화 도구를 사용할 때 위험 제어를 통해 신뢰할 수 있는지를 명확하게 설명하는 것입니다.
Lean is increasingly used to judge natural-language mathematical answers, but its signal is partial: many answers never formalize, and a failed proof may reflect an ill-typed statement or a missing library fact, not a wrong answer. On MATH-500 we show this signal is (i) sharply coverage-dependent, that is the proof-winning answer is correct 96% of the time at high proved coverage but 20% at low, and (ii) sparse and often unfaithful: a 7B autoformalizer proves a class for only 28% of problems, and a manual audit finds only approximately 43% of those proofs faithful. We propose COVCAL, a selector over Lean-trace diagnostics that certifies a finite-sample selective-risk bound on accepted answers or abstains, under two regimes (a conservative Bonferroni bound and a tighter dev-then-cal rule). Feasibility depends on autoformalization coverage: with the 7B formalizer the signal is too sparse and Bonferroni abstains on all 20 bootstrap partitions, whereas a prover-specialized formalizer reaches 79% coverage and flips it to feasible on 17 of 20, accepting approximately 48% of problems at 0.98 accepted accuracy. Since self-consistency alone is already 91% accurate, our contribution is a precise account of when, and with which formalizer, a partial formal signal can be trusted under risk control.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.