LLM이 과학적 발견에 준비되었는가? AI 과학자를 위한 능력 중심 벤치마크
Are LLMs Ready for Scientific Discovery? A Capability-Oriented Benchmark for AI Scientists
기존의 과학 데이터 분석 벤치마크는 LLM을 주로 코드 실행 또는 워크플로우 완료 측면에서 평가하지만, 과학적 분석은 가설 탐색, 통계적 추론, 메커니즘 설명 등 다양한 유형의 과학적 주장을 뒷받침하는 역할을 한다는 점을 간과합니다. 각 유형의 주장에는 서로 다른 가정과 타당성 기준이 존재합니다. 우리는 SDABench라는 벤치마크를 소개하며, 이 벤치마크는 평가를 여섯 가지 능력(기술적, 탐색적, 추론적, 예측적, 인과적, 메커니즘적)을 중심으로 다섯 개 분야(생물학, 화학, 환경, 지리학, 물리학)에 걸쳐 재구성합니다. SDABench는 527개의 실제 데이터 인스턴스(SDA-Real)와 6000개의 합성 데이터 인스턴스(SDA-Synth)로 구성되어 있으며, 각 인스턴스는 객관식 및 자유 서술형 형식으로 제공되며 자동화된 파이프라인을 통해 구축되었습니다. 15개의 대표적인 LLM을 평가한 결과, 모델은 기술적 분석에서는 좋은 성능을 보이지만, 가정 선택, 잠재 과정 모델링 또는 메커니즘 추론이 필요한 작업에서는 성능이 급격히 저하되는 것으로 나타났습니다. 또한, SDABench는 LLM의 오류를 분석하는 다섯 단계 프레임워크를 제공하며, 이를 통해 LLM이 실패하는 지점을 파악할 수 있습니다. 보다 발전된 모델은 관련 범위와 변수를 더 안정적으로 식별하지만, 여전히 적절한 분석 절차를 선택하고, 변수 간 관계를 모델링하고, 타당한 결론을 도출하는 데 어려움을 겪습니다.
Existing benchmarks for scientific data analysis evaluate LLMs primarily on code execution or workflow completion, overlooking that scientific analysis serves to support distinct types of scientific claims: hypothesis exploration, statistical inference, mechanistic explanation, each with different assumptions and validity criteria. We introduce SDABench, a benchmark that reorganizes evaluation around six capabilities (descriptive, exploratory, inferential, predictive, causal, and mechanistic) across five domains (Biology, Chemistry, Environment, Geography, Physics). SDABench comprises 527 real-data instances (SDA-Real) and 6000 synthetic instances (SDA-Synth), each in both multiple-choice and open-ended formats, constructed through an automated pipeline. Evaluating 15 representative LLMs, we find that models handle descriptive analysis well but degrade sharply on tasks requiring assumption selection, latent-process modeling, or mechanistic reasoning. SDABench further provides a five-stage error analysis framework that locates where LLMs fail: more advanced models more reliably identify the relevant scope and variables, but still struggle to select appropriate analytical procedures, model variable relationships, and draw valid conclusions.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.