StatABench: LLM의 통계 분석 능력을 평가하기 위한 데이터셋 및 프레임워크
StatABench: Dataset and Framework for Evaluating Statistical Analysis Capabilities of LLMs
통계 분석은 전문 지식과 도구 활용 능력 모두를 요구하는 광범위하고 복잡한 분야입니다. 기존 연구에서는 대규모 언어 모델(LLM)을 이 분야에서 평가했지만, 현재 벤치마크는 범위와 형식 면에서 제한적입니다. 이러한 격차를 해소하기 위해, LLM의 통계 분석 능력을 체계적으로 평가하도록 설계된 벤치마크인 StatABench (Statistical Analysis Benchmark)를 소개합니다. StatABench는 상호 보완적인 두 가지 구성 요소로 이루어져 있습니다. 첫째, 18가지 통계 주제에 걸쳐 다양한 형식(객관식, 단답형, 의사 결정, 실용적 응용)으로 구성된 404개의 질문을 포함하는 'Stat-Closed'가 있고, 둘째는 전문적인 경쟁에서 가져온 30개의 복잡한 개방형 모델링 과제를 특징으로 하는 'Stat-Open'이 있습니다. LangChain MCP 프레임워크와 다양한 데이터 과학 에이전트를 사용하여 다양한 LLM을 평가하고, 검증된 LLM-as-Judge 프로토콜을 통해 Stat-Open 솔루션을 평가합니다. 실험 결과, GPT-5.1조차도 Stat-Closed에서 68.6%의 정확도를 기록하는 데 그쳤으며, 최고의 오픈 소스 모델은 60.6%를 달성했습니다. Stat-Open에서는 최상위 에이전트 프레임워크가 평균 61.86점을 기록했습니다. 이러한 결과는 현재 LLM과 신뢰할 수 있는 통계 분석 간의 격차를 보여주며, 도구 기반 추론, 방법론적 의사 결정 및 전체적인 통계 모델링에서 여전히 해결해야 할 과제가 있음을 강조합니다.
Statistical analysis is a broad, complex field requiring both domain knowledge and tool proficiency. While prior work has evaluated large language models (LLMs) in this domain, existing benchmarks remain limited in scope and format. To bridge this gap, we introduce StatABench (Statistical AnalysisBenchmark), a benchmark designed to systematically assess LLMs' statistical analysis capabilities. StatABench comprises two complementary components: Stat-Closed, containing 404 questions across 18 statistical topics in multiple formats (multiple-choice, fill-in-the-blank, decision-making, and practical application), and Stat-Open, featuring 30 complex open-ended modeling tasks adapted from professional competitions. We evaluate diverse LLMs using the LangChain MCP framework and multiple data science agents, and assess Stat-Open solutions via a validated LLM-as-Judge protocol. Experiments show that even GPT-5.1 achieves only 68.6% on Stat-Closed, while the best open-source model reaches 60.6%. On Stat-Open, the top agent framework scores 61.86 on average. These results reveal the gap between current LLMs and reliable statistical analysis, highlighting persistent challenges in tool-grounded reasoning, methodological decision-making, and end-to-end statistical modeling.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.