SciR: LLM의 과학적 추론을 위한 제어 가능한 벤치마크
SciR: A Controllable Benchmark for Scientific Reasoning in LLMs
과학적 추론에는 연역, 귀납, 인과적 추론이라는 세 가지 핵심적인 추론 방식이 반복적으로 나타납니다. 현재까지 이러한 방식을 사용하여 LLM을 과학 분야에서 신뢰성 있게 평가하는 것은 어려운 과제입니다. 인간 주석 기반의 과학 벤치마크는 비용이 많이 들고 메커니즘적 진실성을 갖추지 못하며, 인공적인 논리 추론 벤치마크는 실제 과학 문서와 유사하지 않습니다. 본 연구에서는 다중 추론 방식을 통합하고 제어 가능한 과학적 표현을 제공하는 벤치마크인 SciR을 소개합니다. SciR은 세 가지 핵심적인 과학 문제를 기반으로 구축되었습니다. 정형화된 객체(연역 트리, 귀납 규칙 가설, 인과 그래프)에서 작업이 생성되어 검증 가능한 답변을 보장하며, 각 추론 방식에 맞는 전문 분야별 장르를 사용하여 여러 문서로 구성된 과학적 담론으로 표현됩니다. 이러한 구조는 두 가지 난이도 차원을 독립적으로 조정할 수 있도록 합니다. 즉, 추론에 필요한 핵심 정보를 추출하는 데 드는 어려움과, 자체적인 추론 과정의 어려움을 조절할 수 있습니다. 우리는 6개의 모델을 테스트했습니다. 두 가지 난이도 모두 모든 모델에게 영향을 미치며, 그 영향은 복합적으로 나타납니다. 심지어 검증된 솔버에 추론을 위임하는 신경-기호 파이프라인조차도 표현 과정에서 어려움을 겪습니다. 두 가지 난이도 차원은 각 모델의 정보 추출 능력과 추론 능력 간의 관계를 보여줍니다. 예를 들어, deepseek-r1과 같은 추론 모델은 대부분 비추론형 지시 모델보다 추론 능력이 뛰어납니다. SciR은 우리가 알고 있는 한, 추출 및 추론 난이도를 모두 파라미터적으로 제어할 수 있는 최초의 다중 추론 과학 벤치마크입니다.
Three paradigmatic forms of inference recur across scientific reasoning: deduction, induction, and causal abduction. Reliably evaluating LLMs on these in scientific settings is currently out of reach: scientific benchmarks built on human annotations are costly and lack mechanistic ground truth, while synthetic logical-reasoning benchmarks do not resemble real scientific documents. We introduce SciR, a benchmark that combines multi-paradigm reasoning with controllable scientific rendering, anchored on three paradigmatic scientific problems. Tasks are generated from formal objects (deduction tree, inductive rule hypothesis, causal graph) to guarantee verifiable answers, then rendered into multi-document scientific discourse via per-track domain-tuned genres. The construction lets us independently vary two difficulty axes: how hard it is to extract the key information needed for inference, and how hard the principled inference itself is. We test six models. Both axes hurt every model, and their effects compound. The rendering even hurts neurosymbolic pipelines, which hand inference to a verified solver. The two axes yield a per-model extraction-vs-inference profile: for instance, reasoning models like deepseek-r1 mostly surpass non-reasoning instruct models on the inference axis. To our knowledge, SciR is the first multi-paradigm scientific-reasoning benchmark with parametric control on both extraction and inference difficulty.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.