생물학적 프로토콜의 자동 생성 및 실행을 위한 자기 진화형 에이전트 시스템
A Self-Evolving Agentic System for Automated Generation and Execution of Biological Protocols
자율적인 실험실 연구는 단순히 타당한 프로토콜 텍스트 이상을 요구합니다. 생물학적 의도, 정량적 절차, 장치 제약 조건 및 실험 피드백은 프로토콜 및 표준 운영 절차(SOP) 설계부터 코드 작성 및 물리적 실행에 이르기까지 일관성을 유지해야 합니다. 우리는 ProtoPilot이라는 자기 진화형 멀티 에이전트 시스템을 개발했으며, 이 시스템의 전환 과정을 실험 자동화 문제로 테스트하기 위한 전문가 기반 벤치마크 및 평가 프레임워크를 함께 구축했습니다. 이 프레임워크는 98개의 표준 프로토콜, 실험실 전문가 지침, 장치 수준 유효성 검사 기준 및 실제 실험 테스트에서 파생된 294개의 합성 생물학 및 분자 생물학 작업으로 구성됩니다. ProtoPilot은 계층별 검증 가능성, 멀티 에이전트 오케스트레이션 및 실시간 업데이트되는 기술 라이브러리를 통합하여 프로토콜을 생성하고, SOP를 확장하며, SDK 규격에 맞는 코드를 합성하고, 실험실 피드백을 통해 워크플로우를 수정합니다. ProtoPilot은 전문가 선호도 상위 3개 항목 중 하나로 선정될 비율이 90.2%, 전체 프로토콜-코드 변환 성공률이 89.5%였으며, Opentrons 시스템의 성공률은 88.24%였습니다. 이는 OpenTrons-AI의 32.35%에 비해 높은 수치입니다. 실험실 검증을 통해 해석 가능한 결과 데이터, Sanger 시퀀싱으로 확인된 생성물 및 피드백 기반 PCA(주성분 분석) 조립 DNA 타겟이 얻어졌으며, 이는 자율적인 실험을 위한 검증 가능한 경로를 제시합니다. 종합적으로 볼 때, 이러한 결과는 평가 프레임워크가 자율적인 실험실 자동화에 필요한 실행 관련 요구 사항을 반영하고 있으며, ProtoPilot이 프로토콜 및 코드 생성을 검증된 실행 및 피드백 기반 수정으로 변환함으로써 이러한 요구 사항을 충족할 수 있음을 보여줍니다.
Autonomous wet-lab experimentation requires more than plausible protocol text: biological intent, quantitative procedures, device constraints and experimental feedback must remain aligned from protocol and SOP design to code and physical execution. We developed ProtoPilot, a self-evolving multi-agent system, together with an expert-grounded benchmark and evaluation framework for testing this conversion as an experimental automation problem. The framework spans 294 synthetic-biology and molecular-biology tasks derived from 98 gold-standard protocols, wet-lab expert rubrics, device-level validity gates and real experimental tests. ProtoPilot incorporates layer-wise verifiability, multi-agent orchestration and a runtime-updated skill library to generate protocols, expand SOPs, synthesize SDK-compliant code and revise workflows from wet-lab feedback. It achieved a Top@3 expert-preference rate of 90.2%, an overall protocol-to-code gate pass rate of 89.5% and an Opentrons pass rate of 88.24%, compared with 32.35% for OpenTrons-AI. Wet-lab validation produced interpretable readouts, Sanger-confirmed products and feedback-corrected PCA-assembled DNA targets, establishing a verifiable route to autonomous experimentation. Together, these results show that the evaluation framework captures execution-relevant requirements for autonomous wet-lab automation, and that ProtoPilot can meet them by converting protocol and code generation into validated execution and feedback-guided revision.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.