2606.13473v1 Jun 11, 2026 cs.LG

MaxProof: 생성적 검증 강화 학습 및 모집단 수준의 테스트 시간 확장을 통한 수학 증명 확장

MaxProof: Scaling Mathematical Proof with Generative-Verifier RL and Population-Level Test-Time Scaling

Jiacheng Chen
Jiacheng Chen
Citations: 455
h-index: 4
Weiyu Cheng
Weiyu Cheng
Citations: 394
h-index: 4
Zehan Li
Zehan Li
Citations: 5
h-index: 2
Binyan Jiang
Binyan Jiang
Citations: 262
h-index: 4
Han Ding
Han Ding
Citations: 1,242
h-index: 5
Pengyu Zhao
Pengyu Zhao
Citations: 74
h-index: 3
Jingyang Li
Jingyang Li
Citations: 32
h-index: 2
F. Yu
F. Yu
Citations: 15
h-index: 2
Shunkai Zhang
Shunkai Zhang
Citations: 7
h-index: 1
Zhengmao Zhu
Zhengmao Zhu
Citations: 245
h-index: 3
Xinyu Zhang
Xinyu Zhang
Citations: 16
h-index: 2
Yanmohan Wang
Yanmohan Wang
Citations: 0
h-index: 0
Lin Li
Lin Li
Citations: 164
h-index: 2
Tiancheng Qin
Tiancheng Qin
Citations: 56
h-index: 4
Qin Wang
Qin Wang
Citations: 6
h-index: 1
Tianle Li
Tianle Li
Citations: 25
h-index: 3
Jin-Feng Zhu
Jin-Feng Zhu
Citations: 467
h-index: 12
Chenyu Du
Chenyu Du
Citations: 23
h-index: 3
Zijian Song
Zijian Song
Citations: 195
h-index: 3
Jiayuan Song
Jiayuan Song
Citations: 412
h-index: 4
Zhi Zhang
Zhi Zhang
Citations: 141
h-index: 4
Yunan Huang
Yunan Huang
Citations: 376
h-index: 4
Yuntao Cheng
Yuntao Cheng
Citations: 19
h-index: 1

본 논문에서는 MiniMax-M3 시리즈의 경쟁 수준의 수학 증명을 위한 모집단 수준의 테스트 시간 확장 프레임워크인 MaxProof를 제시합니다. M3는 먼저 세 가지 증명 관련 능력을 훈련합니다. 즉, 증명 생성, 증명 검증 및 비판 기반 증명 수정입니다. 이러한 능력은 낮은 오탐율을 갖도록 설계된 심층 방어적 생성 검증기를 사용하여 훈련됩니다. 이렇게 훈련된 기능들은 하나의 통합된 M3 모델로 결합됩니다. 테스트 시간에 MaxProof는 이 모델을 생성기, 검증기, 개선기 및 순위기로 활용하여 후보 증명의 모집단에서 탐색하고, 토너먼트 방식으로 최종 증명을 선택합니다. MaxProof의 테스트 시간 확장을 통해 M3 모델은 IMO 2025에서 35/42점, USAMO 2026에서 36/42점을 달성하여 양 대회 모두 인간 금메달 기준을 초과했습니다.

Original Abstract

We present MaxProof, a population-level test-time scaling framework for competition-level mathematical proof in the MiniMax-M3 series. M3 first trains three proof-oriented capabilities -- proof generation, proof verification, and critique-conditioned proof repair -- using a defense-in-depth generative verifier engineered for low false-positive rate. These capabilities are merged into a single released M3 model. At test time, MaxProof treats the model as a generator, verifier, refiner, and ranker, searches over a population of candidate proofs, and returns one final proof through tournament selection. With MaxProof test-time scaling, the M3 model reaches 35/42 on IMO 2025 and 36/42 on USAMO 2026, exceeding the human gold-medal threshold on both.

1 Citations
0 Influential
6 Altmetric
31.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!