2607.19262v1 Jul 21, 2026 cs.AI

BioSecBench-Surveillance: 병원체 유전체 감시를 위한 인공지능 에이전트 성능 검증을 위한 신뢰성 있는 벤치마크

BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance

Kenny Workman
Kenny Workman
Citations: 5
h-index: 2
Harmon Bhasin
Harmon Bhasin
Citations: 4
h-index: 1
Arjun Banerjee
Arjun Banerjee
Citations: 1
h-index: 1
Kevin Flyangolts
Kevin Flyangolts
Citations: 86
h-index: 5
Dianzhuo Wang
Dianzhuo Wang
Citations: 34
h-index: 4
Evan Seeyave
Evan Seeyave
Citations: 0
h-index: 0
A. Darling
A. Darling
Citations: 106
h-index: 7
Joshua Stallings
Joshua Stallings
Citations: 2
h-index: 1
David Stern
David Stern
Citations: 0
h-index: 0
S. Higdon
S. Higdon
Citations: 346
h-index: 10
C. Duvallet
C. Duvallet
Citations: 24,279
h-index: 26
Bryan Tegomoh
Bryan Tegomoh
Citations: 250
h-index: 3

병원체 유전체 감시 시스템이 확장됨에 따라, 데이터 생성 단계에서 발생하는 병목 현상은 분석 단계로 이동하고 있습니다. 본 논문에서는 BioSecBench-Surveillance이라는 신뢰성 있는 벤치마크를 제시합니다. 이 벤치마크는 인공지능 에이전트가 원시 시퀀스 데이터와 감시 맥락으로부터 올바른 분석 파이프라인을 추론할 수 있는지 여부를 검증하기 위한 100개의 평가 항목으로 구성되어 있습니다. 각 평가에서는 에이전트에게 인간 분석가가 사용할 수 있는 데이터와 맥락만 제공하고, 그 답변의 구조를 결정적으로 평가합니다. 이러한 작업은 분류학적 분류부터 유전자 조작 탐지에 이르기까지 다양한 범주를 포함하며, 다양한 샘플 유형과 시퀀싱 기술을 활용합니다. 16개의 모델-프레임워크 조합에서 총 3,962번의 평가 결과, 가장 뛰어난 성능을 보인 조합은 약 절반 정도의 정확도를 기록했습니다. Opus 4.8 with PI는 83개 평가 항목에서 50.2%의 정확도를 기록했으며, 신뢰 구간은 40.1%에서 60.3%였습니다. 이는 GPT-5.5 with Codex와 동률입니다 (95% 신뢰 구간: 40.8% - 59.6%). 그 뒤를 이어 Opus 4.7 with PI는 49.6%의 정확도를, Sonnet 4.6 with PI는 48.6%의 정확도를 기록했습니다 (각각 95% 신뢰 구간: 40.0% - 59.2%, 38.9% - 58.3%). 에이전트가 올바른 워크플로우를 실행하더라도, 참조 데이터, 임계값, 필터 및 정규화와 같은 요소 선택에서 오류가 발생하는 경우가 있었습니다. BioSecBench-Surveillance은 차세대 발생 시 인공지능 에이전트가 유전체 감시 작업을 신뢰할 수 있는지 여부를 측정하기 위한 표준을 제공합니다.

Original Abstract

As pathogen genomic surveillance scales, the bottleneck is shifting from data generation to analysis. We present BioSecBench-Surveillance, a verifiable benchmark of 100 evaluations testing whether AI agents can infer the right analysis pipeline from raw sequencing data and surveillance context. Each evaluation gives an agent only the data and context a human analyst would have, then grades its structured answer deterministically. The tasks span seven categories, from taxonomic classification to genetic-engineering detection, across diverse sample types and sequencing technologies. Across 3,962 gradable attempts from sixteen model-harness pairs, the strongest configuration cleared only about half. Opus 4.8 with PI led at 50.2 percent, with a 95 percent confidence interval of 40.1 to 60.3 percent across 83 evaluations, tied with GPT-5.5 with Codex at 50.2 percent, with a 95 percent confidence interval of 40.8 to 59.6 percent, followed by Opus 4.7 with PI at 49.6 percent, with a 95 percent confidence interval of 40.0 to 59.2 percent, and Sonnet 4.6 with PI at 48.6 percent, with a 95 percent confidence interval of 38.9 to 58.3 percent. Even when agents invoked the correct workflows, their mistakes came from the choices around them, such as which references, thresholds, filters, and normalization to apply. BioSecBench-Surveillance provides a standard for measuring whether agents can be trusted to perform genomic surveillance when the next outbreak arrives.

0 Citations
0 Influential
13 Altmetric
65.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!