2606.19899v1 Jun 18, 2026 cs.CY

AI 에이전트의 생물학적 능력 및 위험 측정

Measuring Biological Capabilities and Risks of AI Agents

Patricia Paskov
Patricia Paskov
Citations: 62
h-index: 4
Jeffrey Lee
Jeffrey Lee
Citations: 1
h-index: 1
Kyle R. Brady
Kyle R. Brady
Citations: 62
h-index: 3
Alyssa Worland
Alyssa Worland
Citations: 124
h-index: 6

본 논문은 급속하게 부상하고 있는 정책 과제인 AI 과학자 또는 자율적으로 또는 협력하여 다단계 과학적 작업을 수행할 수 있는 에이전트형 AI 시스템의 생물학적 능력과 위험에 대한 신뢰성 있는 증거를 어떻게 생성하고 해석할 것인가에 대해 다룬다. 이러한 시스템이 실제 연구 워크플로우에 진입함에 따라, 정책 결정자들은 종종 암묵적이거나 문서화되지 않은 기본 설계 선택에 그 의미가 좌우되는 평가 결과에 직면하게 된다. 본 논문은 AI 기반 생물학적 위험에 대한 현재 증거를 종합하고, 해석에 민감하지만 유망한 도구인 생물학적 에이전트 평가를 소개한다. 핵심적인 기여는 실제 경험을 바탕으로 한 실용적인 고려 사항들을 제시하는데, 이는 평가의 정의, 설계, 실행, 점수 산정 및 문서화 방식이 결과가 나타내는 위험에 대해 어떤 의미를 가지는지에 대한 이해를 돕는다. 본 분석은 정책 입안자들이 생물학적 평가 결과를 적절한 주의를 기울여 해석하도록 돕고, 공공 및 민간 자금 지원 기관들이 AI-생물학 평가 연구에 고효율 투자를 할 수 있도록 안내하며, 바이오 보안 전문가들이 새롭게 등장하는 AI 시스템을 평가하는 데 기여하는 것을 목표로 한다. 또한, 선도적인 AI 연구소, AI 제공 업체, 과학 기관 및 제3자 평가 기관에서 에이전트형 평가를 설계하거나 수행하는 연구자들을 위한 자료가 될 것이다.

Original Abstract

This paper addresses a rapidly emerging policy challenge: how to generate and interpret credible evidence about the biological capabilities and risks of AI scientists, or agentic AI systems capable of autonomously or collaboratively performing multi-step scientific tasks. As these systems enter real research workflows, decision-makers increasingly face evaluation results whose meaning depends on underlying design choices that are often implicit or under-documented. We synthesize current evidence on AI-enabled biological risks and introduce biological agentic evaluations as a promising, but interpretation-sensitive, tool for assessing these systems. Our central contribution is a set of practical, experience-grounded considerations -- drawing from our own evaluations -- that show how choices around defining, designing, running, scoring, and documenting evaluations materially shape what results do and do not imply about risk. The analysis is intended to help policymakers interpret biological evaluation outputs with appropriate caution; guide public and private funders toward high-leverage investments in AI-biology evaluation research; and support biosecurity practitioners assessing emerging AI systems. A secondary audience includes researchers designing or conducting agentic evaluations within frontier AI labs, AI providers, scientific institutions, and third-party evaluation organizations.

4 Citations
0 Influential
3 Altmetric
19.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!