대규모 언어 모델을 활용한 다양한 과학적 가설 탐색 연구
Towards Diverse Scientific Hypothesis Search with Large Language Models
대규모 언어 모델(LLM)은 최근 과학적 발견을 가속화하는 데 중요한 역할을 하고 있으며, 특히 유효한 과학적 가설 생성과 같은 고급 작업에 활용되고 있습니다. 그러나 많은 과학적 발견 과정에서 가장 좋은 하나의 가설을 찾는 것만이 목표가 아니기 때문에, 검증 과정이 불확실하고 비용이 많이 들 수 있으며, 연구자들은 최상의 결과를 얻기 위해 다양한 고품질의 대체 가설 세트를 보유하는 것이 중요합니다. 일반적으로 사용되는 진화적 탐색 방법은 가설 생성 과정에서 최적화에 지나치게 집중하는 경향이 있으며, 이로 인해 검색 과정에서의 선택 압력이 다양성을 감소시키는 문제를 야기합니다. 이러한 한계점을 극복하기 위해, 우리는 가설 탐색을 효율적으로 다양한 고품질 가설을 생성하는 샘플링 문제로 정의하고, 제한된 검증 예산 내에서 이를 달성하고자 합니다. 이러한 관점에서, 우리는 고전적인 병렬 템퍼링 알고리즘에서 영감을 받은 진화적 프레임워크인 ours 를 제안합니다. 이 프레임워크는 여러 온도 수준에서 가설을 탐색하고, 수렴을 방해하지 않으면서 원활한 정보 교환을 통해 탐색 성능을 향상시킵니다. 분자 발견, 방정식 발견 및 알고리즘 발견과 같은 다양한 분야에서 우리의 접근 방식은 동일한 검증 예산 내에서 가설의 품질과 다양성을 모두 개선하며, 더 비싼 후속 계산 검증 과정에서도 안정적인 결과를 제공합니다.
Large language models (LLMs) are on the rise for accelerating scientific discovery, most recently in advanced tasks such as generating valid scientific hypotheses. Yet in many discovery settings, the goal is not to identify a single best hypothesis since validation can be noisy and expensive, and scientists benefit from a set of high-quality alternative hypotheses that hedge against downstream uncertainty for the best solutions. Nevertheless, commonly used evolutionary search recipes tend to prioritize optimization over exploration in hypothesis generation, and the resulting selection pressure during the search process leads to diversity collapse. Motivated by these limitations, we formulate hypothesis search as a sampling problem, where the objective is to efficiently produce diverse, high-quality hypotheses under a fixed validation budget. Building on this perspective, we propose \ours, an evolutionary framework inspired by the classical parallel tempering algorithm that searches hypotheses at multiple temperature levels and enables principled information exchange across temperatures to improve exploration without disrupting convergence. Across domains including molecular discovery, equation discovery, and algorithm discovery, our approach consistently improves both hypothesis quality and diversity under the same validation budget, and produces candidates that remain robust under more expensive downstream computational validations.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.