2505.20321v6 May 23, 2025 cs.CL

BiomedSQL: 생체 의학 지식 베이스 기반 과학적 추론을 위한 텍스트-to-SQL

BiomedSQL: Text-to-SQL for Scientific Reasoning on Biomedical Knowledge Bases

Daniel Khashabi
Daniel Khashabi
Citations: 353
h-index: 9
Mathew J. Koretsky
Mathew J. Koretsky
Citations: 442
h-index: 11
Maya Willey
Maya Willey
Citations: 34
h-index: 4
Adi Asija
Adi Asija
Citations: 25
h-index: 3
Owen Bianchi
Owen Bianchi
Citations: 26
h-index: 4
Chelsea X. Alvarado
Chelsea X. Alvarado
Citations: 1,378
h-index: 10
Tanay Nayak
Tanay Nayak
Citations: 13
h-index: 2
Nicole Kuznetsov
Nicole Kuznetsov
Citations: 115
h-index: 6
Sungwon Kim
Sungwon Kim
Citations: 32
h-index: 3
M. Nalls
M. Nalls
Citations: 87,021
h-index: 134
F. Faghri
F. Faghri
Citations: 7,466
h-index: 34

생명 의학 연구자들은 점점 더 복잡한 분석 작업을 위해 대규모 구조화된 데이터베이스에 의존하고 있습니다. 그러나 현재의 텍스트-to-SQL 시스템은 종종 질적인 과학 질문을 실행 가능한 SQL로 변환하는 데 어려움을 겪으며, 특히 암묵적인 도메인 추론이 필요한 경우 더욱 그렇습니다. 본 논문에서는 실제 생체 의학 지식 베이스에 대한 텍스트-to-SQL 생성에서 과학적 추론을 평가하도록 특별히 설계된 최초의 벤치마크인 BiomedSQL을 소개합니다. BiomedSQL은 템플릿 기반으로 생성되고 BigQuery 데이터베이스에 통합된 유전자-질병 연관, 오믹스 데이터를 통한 인과 관계 추론 및 약물 승인 기록을 포함하는 조화된 데이터베이스를 기반으로 하는 68,000개의 질문/SQL 쿼리/답변 세트로 구성됩니다. 각 질문은 모델이 유전체 전체 유의성 임계값, 효과 방향 또는 시험 단계 필터링과 같은 도메인별 기준을 추론하도록 요구하며, 단순히 구문 분석 번역에만 의존하지 않습니다. 우리는 다양한 오픈 소스 및 상용 LLM을 프롬프트 전략 및 상호 작용 방식 측면에서 평가했습니다. 그 결과 상당한 성능 격차가 나타났습니다. 기본 프롬프팅 하에서 Gemini-3-Pro는 58.1%의 실행 정확도를 달성하는 반면, 당사 맞춤형 다단계 에이전트인 BMSQL은 62.6%에 도달했으며, 이는 전문가 기준인 90.0%보다 훨씬 낮습니다. BiomedSQL은 구조화된 생체 의학 지식 베이스에 대한 강력한 추론을 통해 과학적 발견을 지원하는 텍스트-to-SQL 시스템을 발전시키는 새로운 기반을 제공합니다. BiomedSQL 벤치마크 및 코드는 https://datatecnica.github.io/biomedbench-suite/biomedsql 에서 공개적으로 이용할 수 있습니다.

Original Abstract

Biomedical researchers increasingly rely on large-scale structured databases for complex analytical tasks. However, current text-to-SQL systems often struggle to map qualitative scientific questions into executable SQL, particularly when implicit domain reasoning is required. We introduce BiomedSQL, the first benchmark explicitly designed to evaluate scientific reasoning in text-to-SQL generation over a real-world biomedical knowledge base. BiomedSQL comprises 68,000 question/SQL query/answer triples generated from templates and grounded in a harmonized BigQuery database that integrates gene-disease associations, causal inference from omics data, and drug approval records. Each question requires models to infer domain-specific criteria, such as genome-wide significance thresholds, effect directionality, or trial phase filtering, rather than rely on syntactic translation alone. We evaluate a range of open- and closed-source LLMs across prompting strategies and interaction paradigms. Our results reveal a substantial performance gap: Gemini-3-Pro achieves 58.1% execution accuracy under baseline prompting, while our custom multi-step agent, BMSQL, reaches 62.6%, both well below the expert baseline of 90.0%. BiomedSQL provides a new foundation for advancing text-to-SQL systems that support scientific discovery through robust reasoning over structured biomedical knowledge bases. The BiomedSQL benchmark and codebase are publicly available at https://datatecnica.github.io/biomedbench-suite/biomedsql.

7 Citations
0 Influential
30 Altmetric
157.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!