2606.11150v1 Jun 09, 2026 cs.AI

ABC-Bench: 생물 보안을 위한 능동적 생체 능력 평가 기준

ABC-Bench: An Agentic Bio-Capabilities Benchmark for Biosecurity

Alexander Kleinman
Alexander Kleinman
Citations: 3
h-index: 1
Bryce Cai
Bryce Cai
Citations: 3
h-index: 1
A. B. Liu
A. B. Liu
Citations: 47
h-index: 4
S. Nedungadi
S. Nedungadi
Citations: 39
h-index: 2
Seth Donoughe
Seth Donoughe
Citations: 48
h-index: 2
Harmon Bhasin
Harmon Bhasin
Citations: 4
h-index: 1

대규모 언어 모델(LLM)은 문헌 종합부터 실험 데이터 해석에 이르기까지 생물학 연구와 관련된 기능을 빠르게 습득하고 있습니다. 또한, LLM 에이전트는 이전에는 경험 있는 인간 생물학자가 수행해야 했던 인공 생물학 작업을 수행할 수 있게 되었습니다. 이러한 새로운 AI 기능은 과학적 발견과 생의학 발전에 새로운 기회를 제공하지만, 동시에 생물 보안 위험의 지형을 변화시킵니다. 이에 대응하기 위해, 우리는 능동적 생체 능력 평가 기준(ABC-Bench)을 소개합니다. ABC-Bench는 에이전트의 생물 보안 관련 능력을 측정하는 일련의 작업으로 구성되어 있습니다. ABC-Bench는 LLM 에이전트를 양성적인 작업과 이중 용도 생물학 작업 모두에서 평가합니다. 여기에는 액체 처리 로봇을 작동하는 코드를 작성하고, 시험관 내 조립을 위한 DNA 단편을 설계하며, DNA 합성 검사를 회피하는 작업이 포함됩니다. 이러한 작업은 생물학과 소프트웨어 전문 지식을 결합해야 합니다. 테스트된 모든 LLM 에이전트는 세 가지 작업 모두에서 중간 수준의 전문가 인간 기준선보다 뛰어난 성능을 보였습니다. 에이전트는 공개된 지식과 잘 문서화된 프로토콜에 기반한 작업에서 높은 성과를 보였지만, 새로운 생물 정보학적 추론이 필요한 작업에서는 상대적으로 낮은 성과를 보였습니다. 세 가지 실험실 검증 실험에서 OpenAI의 o4-mini-high 모델은 OpenTrons 액체 처리 로봇에서 실행했을 때 예상되는 서열을 가진 DNA를 성공적으로 조립하는 스크립트를 생성했습니다.

Original Abstract

Large language models (LLMs) are rapidly acquiring capabilities relevant to biological research, from literature synthesis to interpretation of experimental data. Increasingly, LLM agents can also perform in silico biology tasks that previously required experienced human biologists. These emerging AI capabilities offer new opportunities for scientific discovery and biomedical advances, but they also shift the landscape of biosecurity risks. To address this, we introduce the Agentic Bio-Capabilities Benchmark (ABC-Bench), a suite of tasks to measure agentic biosecurity-relevant capabilities. ABC-Bench evaluates LLM agents on both benign and dual-use biology tasks: writing code to operate liquid handling robots, designing DNA fragments for in vitro assembly, and evading DNA synthesis screening. These tasks require a combination of biology and software expertise. All tested LLM agents outperformed the median expert human baseliner on all three tasks. Agents performed highly on tasks drawing on published knowledge and well-documented protocols, and more weakly on a task requiring novel bioinformatics reasoning. In three wet-lab validation experiments, we found that OpenAI's o4-mini-high produced scripts that, when run on an OpenTrons liquid handling robot, successfully assembled DNA with expected sequences.

5 Citations
0 Influential
2 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!