InfiniteScienceGym: 과학적 분석을 위한 무한하고 절차적으로 생성된 벤치마크
InfiniteScienceGym: An Unbounded, Procedurally-Generated Benchmark for Scientific Analysis
대규모 언어 모델은 과학적 조력자로서 등장하고 있지만, 경험적 데이터를 기반으로 추론하는 능력을 평가하는 것은 여전히 어려운 과제입니다. 기존 연구 및 인간 주석을 기반으로 한 벤치마크는 출판 편향, 알려진 지식 편향, 레이블 노이즈, 그리고 상당한 저장 공간 요구 사항과 같은 문제점을 안고 있습니다. 본 연구에서는 절차적으로 생성된 벤치마크인 InfiniteScienceGym을 소개합니다. InfiniteScienceGym은 과학적 데이터 저장소를 생성하고, 검증 가능한 질의응답(QA) 작업을 결합합니다. 시드 값을 기반으로, 시뮬레이터는 현실적인 디렉토리 구조, 파일, 테이블 데이터로 구성된 독립적인 저장소를 결정적으로 생성합니다. 또한, 특권 QA 생성기는 정답을 정확하게 포함하는 답변 가능한 질문과 답변 불가능한 질문을 생성합니다. 이를 통해 대규모 정적 데이터 세트를 배포하지 않고도, 증거 기반 추론, 거부(abstention), 그리고 도구 활용 분석을 제어된 환경에서 평가할 수 있습니다. InfiniteScienceGym은 기존의 과학적 벤치마크를 보완하며, 공개된 데이터 세트만으로는 평가하기 어려운 약점과 오류 패턴을 목표로 합니다. 독점 모델과 공개 모델을 모두 평가한 결과, 어떤 모델도 전반적으로 45% 이상의 정확도를 달성하지 못했으며, 답변 불가능한 질문을 식별하는 능력은 여전히 주요한 약점임을 확인했습니다. 또한, 더 강력한 모델은 단순히 더 많은 토큰을 사용하는 것보다 도구를 더 효과적으로 활용하는 경향이 있다는 것을 발견했습니다.
Large language models are emerging as scientific assistants, but evaluating their ability to reason from empirical data remains challenging. Benchmarks derived from published studies and human annotations inherit publication bias, known-knowledge bias, label noise, and substantial storage requirements. We present InfiniteScienceGym, a procedurally generated benchmark of scientific repositories paired with a verifiable question-answering task. From a seed, the simulator deterministically generates a self-contained repository with realistic directory structure, files, and tabular data, and a privileged QA generator produces both answerable and unanswerable questions with exact ground truth. This makes it possible to evaluate evidence-grounded reasoning, abstention, and tool-mediated analysis in a controlled setting without distributing a large static corpus. InfiniteScienceGym complements real scientific benchmarks by targeting blind spots and failure modes that are hard to evaluate using published datasets alone. Evaluating both proprietary and open-weight models, we find that none achieve more than 50% accuracy overall, that recognizing unanswerable questions remains a major weakness, and that stronger models tend to use tools more effectively rather than simply consuming more tokens.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.