SABER-Math: 수학 분야 정보 검색 평가를 위한 자동화된 벤치마크
SABER-Math: Automated Benchmark for Information Retrieval Evaluation in Mathematics
지능형 AI 시스템이 더욱 복잡한 수학적 문제를 해결하면서, 문제 데이터베이스, 정리 라이브러리 및 교육 자료를 검색하기 위해 정보 검색(IR)에 점점 더 의존하고 있습니다. 그러나 적절한 검색 엔진을 선택하는 것은 여전히 어렵습니다. 왜냐하면 검색 엔진의 성능이 최종 결과에 미치는 영향을 직접적으로 분리하여 평가하기가 불가능하기 때문입니다. 반면, 기존의 검색 엔진 특화 벤치마크는 종종 미묘한 수학적 관련성을 제대로 반영하지 못하여 관련 문서에 불이익을 줄 수 있습니다. 우리는 이러한 격차를 해결하기 위해 전문가의 주석 없이 수학 분야 IR을 평가하는 최초의 완전 자동화된 벤치마크인 SABER-Math를 소개합니다. SABER-Math는 283,000개의 고등학교 수준 수학 문제와 해법을 기반으로 구축되었으며, 세 단계를 거쳐 도전적인 재순위화 작업을 수행합니다. (i) 먼저, LLM(Large Language Model)이 각 문제에 대한 간결한 해법 요약 및 수학적 주제를 추출합니다. (ii) 다음으로, 온톨로지 토픽 기반 및 어휘-해법 요약 기반 유사성을 활용하여 쿼리당 관련 문서를 검색합니다. (iii) 마지막으로, 스위스식 LLM 선호도 토너먼트를 통해 문서에 대한 미세한 관련성 점수를 생성합니다. 우리는 어휘 기반 검색 엔진, 특수 수학 검색 시스템 및 최신 임베딩 모델을 평가했습니다. 그 결과, 현대적인 임베딩 모델은 기존 방식 및 수학 전용 기준 성능보다 훨씬 우수한 성능을 보이지만, 대수학 및 미적분학과 같이 기호 중심 영역에서는 가장 강력한 시스템조차 어려움을 겪는다는 것을 확인했습니다. 또한, MTEB과 같은 범용 IR 벤치마크가 최신 임베딩 모델의 수학 분야 성능을 안정적으로 예측하지 못한다는 점을 보여주었습니다. 이는 수학 전용 검색 벤치마크의 필요성을 강조합니다.
As agentic AI systems tackle more complex mathematical tasks, they increasingly rely on information retrieval (IR) to search problem databases, theorem libraries, and educational resources. However, choosing the right retriever remains difficult, as it is infeasible to directly isolate its effect on downstream performance. On the other hand, existing retrieval-specific benchmarks often fail to capture fine-grained mathematical relevance, penalizing relevant documents. We address this gap by introducing SABER-Math, the first fully automated benchmark for evaluating mathematical IR without expert annotation. Starting from 283K high-school-level math problems with solutions, SABER-Math builds challenging reranking tasks in three steps: (i) first, LLMs extract concise solution summaries and mathematical topics for each problem; (ii) then, per-query relevant documents are discovered using ontology topic-based and lexical solutions-summary-based similarities, and (iii) finally, a Swiss-style LLM preference tournament produces fine-grained relevance ratings for the documents. We evaluate lexical retrievers, specialized mathematical retrieval systems, and recent embedding models. We find that while modern embedding models substantially outperform classical and math-specific baselines, even the strongest systems struggle in symbol-heavy domains like Algebra and Calculus. Importantly, we show that general-purpose IR benchmarks such as MTEB do not reliably predict mathematical performance, especially for recent embedding models, highlighting the need for math-specific retrieval benchmarks.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.