MaxProof: 생성적 검증 강화 학습 및 모집단 수준의 테스트 시간 확장을 통한 수학 증명 확장
MaxProof: Scaling Mathematical Proof with Generative-Verifier RL and Population-Level Test-Time Scaling
본 논문에서는 MiniMax-M3 시리즈의 경쟁 수준의 수학 증명을 위한 모집단 수준의 테스트 시간 확장 프레임워크인 MaxProof를 제시합니다. M3는 먼저 세 가지 증명 관련 능력을 훈련합니다. 즉, 증명 생성, 증명 검증 및 비판 기반 증명 수정입니다. 이러한 능력은 낮은 오탐율을 갖도록 설계된 심층 방어적 생성 검증기를 사용하여 훈련됩니다. 이렇게 훈련된 기능들은 하나의 통합된 M3 모델로 결합됩니다. 테스트 시간에 MaxProof는 이 모델을 생성기, 검증기, 개선기 및 순위기로 활용하여 후보 증명의 모집단에서 탐색하고, 토너먼트 방식으로 최종 증명을 선택합니다. MaxProof의 테스트 시간 확장을 통해 M3 모델은 IMO 2025에서 35/42점, USAMO 2026에서 36/42점을 달성하여 양 대회 모두 인간 금메달 기준을 초과했습니다.
We present MaxProof, a population-level test-time scaling framework for competition-level mathematical proof in the MiniMax-M3 series. M3 first trains three proof-oriented capabilities -- proof generation, proof verification, and critique-conditioned proof repair -- using a defense-in-depth generative verifier engineered for low false-positive rate. These capabilities are merged into a single released M3 model. At test time, MaxProof treats the model as a generator, verifier, refiner, and ranker, searches over a population of candidate proofs, and returns one final proof through tournament selection. With MaxProof test-time scaling, the M3 model reaches 35/42 on IMO 2025 and 36/42 on USAMO 2026, exceeding the human gold-medal threshold on both.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.