ComBench: 올림피아드 수준의 조합론 문제 해결을 위한 엄격한 증명 추론 및 구성적 실현 능력 평가 벤치마크
ComBench: A Benchmark for Rigorous Proof Reasoning and Constructive Realization in Olympiad-Level Combinatorics
조합론은 올림피아드 수준의 수학 문제 해결에 필수적인 요소이며, 깊이 있는 이산적 추론, 창의적인 구성, 그리고 엄밀한 구조적 통찰력을 요구합니다. 최근 연구 결과에 따르면, 현재 가장 뛰어난 최첨단 모델조차도 올림피아드 조합론 문제에서 일관된 성능을 보이지 않으며, 이는 수학적 창의성 추론 능력의 격차를 보여줍니다. 본 논문에서는 대규모 언어 모델의 조합론적 추론 능력을 평가하고 진단하기 위한 올림피아드 수준의 조합론 벤치마크인 ComBench를 소개합니다. ComBench는 100개의 인간이 직접 검토한 경쟁 수준의 문제들로 구성되어 있으며, 두 가지 상호 보완적인 환경을 중심으로 분류됩니다. 첫째는 엄격한 수학적 논리를 주로 요구하는 분석 중심 문제이고, 둘째는 정확성 입증 외에도 명시적인 구성을 필요로 하는 구성 중심 문제입니다. 평가 프로토콜은 채점 기준에 따른 증명 평가와 결정론적인 구성 검증을 결합하여, 증명의 품질과 구성의 타당성이 일치하지 않는 사례를 밝혀냅니다. 최첨단 오픈 소스 및 폐쇄 소스 모델에 대한 실험 결과는 ComBench가 아직 충분히 활용되지 않았음을 보여줍니다. 가장 뛰어난 모델은 전체 평균 65.4%, 상위 4개 답변 중 최고 성능을 기준으로 75.3%의 정확도를 보였습니다. 또한, 엄격한 증명 추론과 구성적 실현은 서로 다른 능력이라는 것을 확인했습니다. Kimi-K2.6은 GPT-5.5보다 분석 중심 증명 평가에서는 낮은 점수를 받았지만, 구성 중심 상위 4개 답변 기준으로 더 높은 성능을 보였습니다. 마지막으로, 존재 및 구성 문제는 대표적인 최첨단 모델 전반에 걸쳐 지속적으로 가장 어려운 문제로 나타났습니다.
Combinatorics is central to Olympiad-level mathematical problem solving, requiring deep discrete reasoning, creative constructions, and rigorous structural insight. Recent evidence suggests that even today's strongest frontier models remain uneven on Olympiad combinatorics, revealing a gap in creative mathematical reasoning. We introduce ComBench, an Olympiad-level combinatorics benchmark for evaluating and diagnosing the combinatorial reasoning capabilities of large language models. ComBench contains 100 human-annotated competition-level problems organized around two complementary settings: analysis-centric problems, which primarily require rigorous mathematical arguments, and construction-centric problems, which require explicit constructions in addition to correctness justifications. The evaluation protocol combines rubric-guided proof grading with deterministic construction verification, exposing cases where proof quality and construction validity diverge. Experiments on frontier open- and closed-source models show that ComBench is far from saturated: the strongest model reaches 65.4% overall Avg. and 75.3% overall Best@4. We further find that Rigorous Proof Reasoning and Constructive Realization are distinct capabilities: Kimi-K2.6 trails GPT-5.5 on analysis-centric proof grading but surpasses it on construction-centric Best@4, while Existence and Construction problems remain consistently hardest across representative frontier models.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.