BioSecBench-Surveillance: 병원체 유전체 감시를 위한 인공지능 에이전트 성능 검증을 위한 신뢰성 있는 벤치마크
BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance
병원체 유전체 감시 시스템이 확장됨에 따라, 데이터 생성 단계에서 발생하는 병목 현상은 분석 단계로 이동하고 있습니다. 본 논문에서는 BioSecBench-Surveillance이라는 신뢰성 있는 벤치마크를 제시합니다. 이 벤치마크는 인공지능 에이전트가 원시 시퀀스 데이터와 감시 맥락으로부터 올바른 분석 파이프라인을 추론할 수 있는지 여부를 검증하기 위한 100개의 평가 항목으로 구성되어 있습니다. 각 평가에서는 에이전트에게 인간 분석가가 사용할 수 있는 데이터와 맥락만 제공하고, 그 답변의 구조를 결정적으로 평가합니다. 이러한 작업은 분류학적 분류부터 유전자 조작 탐지에 이르기까지 다양한 범주를 포함하며, 다양한 샘플 유형과 시퀀싱 기술을 활용합니다. 16개의 모델-프레임워크 조합에서 총 3,962번의 평가 결과, 가장 뛰어난 성능을 보인 조합은 약 절반 정도의 정확도를 기록했습니다. Opus 4.8 with PI는 83개 평가 항목에서 50.2%의 정확도를 기록했으며, 신뢰 구간은 40.1%에서 60.3%였습니다. 이는 GPT-5.5 with Codex와 동률입니다 (95% 신뢰 구간: 40.8% - 59.6%). 그 뒤를 이어 Opus 4.7 with PI는 49.6%의 정확도를, Sonnet 4.6 with PI는 48.6%의 정확도를 기록했습니다 (각각 95% 신뢰 구간: 40.0% - 59.2%, 38.9% - 58.3%). 에이전트가 올바른 워크플로우를 실행하더라도, 참조 데이터, 임계값, 필터 및 정규화와 같은 요소 선택에서 오류가 발생하는 경우가 있었습니다. BioSecBench-Surveillance은 차세대 발생 시 인공지능 에이전트가 유전체 감시 작업을 신뢰할 수 있는지 여부를 측정하기 위한 표준을 제공합니다.
As pathogen genomic surveillance scales, the bottleneck is shifting from data generation to analysis. We present BioSecBench-Surveillance, a verifiable benchmark of 100 evaluations testing whether AI agents can infer the right analysis pipeline from raw sequencing data and surveillance context. Each evaluation gives an agent only the data and context a human analyst would have, then grades its structured answer deterministically. The tasks span seven categories, from taxonomic classification to genetic-engineering detection, across diverse sample types and sequencing technologies. Across 3,962 gradable attempts from sixteen model-harness pairs, the strongest configuration cleared only about half. Opus 4.8 with PI led at 50.2 percent, with a 95 percent confidence interval of 40.1 to 60.3 percent across 83 evaluations, tied with GPT-5.5 with Codex at 50.2 percent, with a 95 percent confidence interval of 40.8 to 59.6 percent, followed by Opus 4.7 with PI at 49.6 percent, with a 95 percent confidence interval of 40.0 to 59.2 percent, and Sonnet 4.6 with PI at 48.6 percent, with a 95 percent confidence interval of 38.9 to 58.3 percent. Even when agents invoked the correct workflows, their mistakes came from the choices around them, such as which references, thresholds, filters, and normalization to apply. BioSecBench-Surveillance provides a standard for measuring whether agents can be trusted to perform genomic surveillance when the next outbreak arrives.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.