2605.30284v1 May 28, 2026 cs.AI

ProjectionBench: 점진적인 정보 공개 하에서 LLM의 과학적 가설 생성 능력 평가

ProjectionBench: Evaluating Scientific Hypothesis Generation in LLMs Under Progressive Information Disclosure

M. J. Buehler
M. J. Buehler
Citations: 20
h-index: 3
A. Lew
A. Lew
Citations: 0
h-index: 0
Y. Cao
Y. Cao
Citations: 0
h-index: 0

과학적 발견은 본질적으로 창의적이고 불확실한 과정이며, 알려진 지식을 단순 회상하는 것을 넘어선 추론 능력을 요구합니다. 많은 벤치마크가 다중 정보 검색을 통해 LLM의 심층 연구 수행 능력을 평가하기 위해 제안되었지만, 진정한 과학적 발견에 필수적인 혁신적인 추론 능력은 아직 대부분 검증되지 않았습니다. 본 논문에서는 과학적 발견 및 추론 능력을 평가하는 벤치마크 프레임워크를 소개하며, 이는 원시 문제에서부터 고전적인 귀무 가설 검정에 이르기까지 점진적으로 구성됩니다. 저희의 프레임워크에서 모델은 처음에는 최근 연구 논문의 주제와 연구 질문만을 입력받고, 기술적인 세부 사항이 점진적으로 공개됩니다. 정보가 단계별로 제공될 때마다, 모델은 연구 질문에 대한 가설을 생성하는 과제를 수행하며, 이는 원본 논문의 결론과 비교되고, 구성 요소인 주장의 의미적 유사성을 통해 자동으로 평가됩니다. 이러한 단계별 의미적 일관성 평가는 모델의 혁신적인 능력(최소한의 정보 환경에서)부터 근거 기반 추론 능력(전체 실험 세부 사항 하에서)까지를 평가할 수 있도록 하며, 이는 LLM을 과학적 발견에 활용하는 데 있어 매우 중요합니다. 저희 프레임워크는 LLM의 과학적 추론 및 발견 능력을 체계적으로 평가하기 위한 기반을 제공하며, 차세대 AI 과학자/연구 보조 시스템 개발에 필수적인 역할을 합니다. 특히, 본 연구에서는 GPT-5, GPT-5.4, Gemini 2.5 pro, 그리고 Gemini 3.1 pro preview 모델을 사용하여 생리 활성 재료, 기계적 재료 및 나노 재료 분야의 45개 논문에 대한 평가를 수행했습니다. 결과적으로 예상대로 GPT-5.4와 Gemini 3.1 pro는 이전 세대 모델보다 우수한 성능을 보였으며, 특히 GPT-5.4는 최소한의 맥락에서도 0.7의 F1 점수로 원본 결론과 일치하는 결과를 나타냈습니다.

Original Abstract

Scientific discovery is an inherently creative and uncertain process, requiring reasoning beyond the recall of known knowledge. While many benchmarks have been proposed to evaluate large language model (LLM) performance on deep research tasks via multi-hop retrieval, their innovative reasoning abilities essential for true scientific discovery remain largely untested. We introduce a benchmark framework for evaluating model performance in scientific discovery and reasoning, building up from a raw problem to the classical null hypothesis test. In our framework, models initially receive only the topic and research question from a recent paper, with technical details progressively revealed. At each stage of information disclosure, the model is tasked with generating hypotheses that address the research question, which is compared with the conclusions from the original paper and evaluated via automated semantic similarity of constituent atomic claims. This progressive evaluation of semantic divergence from ground-truth conclusions enables assessment of a model's innovativeness (under minimal information) to grounded reasoning capabilities (under full experimental details), both critical for using LLMs for scientific discovery purposes. Our framework provides a foundation for systematically evaluating scientific reasoning and discovery capabilities in LLMs, crucial for advancing the development of next-generation AI scientist/co-scientist systems. Specifically, here we evaluate GPT-5, GPT-5.4, Gemini 2.5 pro, and Gemini 3.1 pro preview across 45 papers spanning bioactive materials, mechanical materials, and nanomaterials. We find that GPT-5.4 and Gemini 3.1 pro outperform their previous generation counterparts as expected, and GPT-5.4 in particular maintains 0.7 F1 score alignment with ground truth conclusions even under minimal context.

0 Citations
0 Influential
1.5 Altmetric
7.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!