랭크 앤 리즌: 다중 에이전트 협업을 통한 제로샷 단백질 변이 예측 가속화
Rank-and-Reason: Multi-Agent Collaboration Accelerates Zero-Shot Protein Mutation Prediction
제로샷 변이 예측은 제한적인 자원을 가진 단백질 공학에 매우 중요하지만, 기존의 단백질 언어 모델(PLM)은 종종 통계적으로 신뢰할 만한 결과를 제공하면서도 기본적인 생물리학적 제약을 무시하는 경향이 있습니다. 현재, 실험 검증을 위한 후보 물질을 선택하는 과정은 PLM의 출력 결과를 수동으로 검토하는 방식으로 이루어지며, 이는 비효율적이고 주관적이며, 해당 분야의 전문 지식에 크게 의존합니다. 이러한 문제를 해결하기 위해, 우리는 Rank-and-Reason (VenusRAR)이라는 두 단계로 구성된 에이전트 기반 프레임워크를 제안합니다. 이 프레임워크는 워크플로우를 자동화하고, 실험적 성공 가능성을 극대화하는 것을 목표로 합니다. Rank 단계에서는, 계산 전문가와 가상 생물학자가 문맥에 맞는 다중 모드 앙상블을 사용하여 ProteinGym 데이터셋에서 새로운 스피어만 상관관계 기록인 0.551 (vs. 0.518)을 달성합니다. Reason 단계에서는, 에이전트 기반 전문가 패널이 연쇄적 추론(chain-of-thought reasoning)을 사용하여 후보 물질을 기하학적 및 구조적 제약 조건에 맞게 검토하며, ProteinGym-DMS99 데이터셋에서 상위 5개 정확도(Top-5 Hit Rate)를 최대 367%까지 향상시킵니다. Cas12i3 뉴클라아제에 대한 실험 검증은 이 프레임워크의 효능을 더욱 확인시켜 주며, 46.7%의 긍정적인 결과를 얻었으며, 4.23배 및 5.05배의 활성 개선을 보이는 두 가지 새로운 변이체를 식별했습니다. 코드 및 데이터셋은 GitHub (https://github.com/ai4protein/VenusRAR/)에서 공개됩니다.
Zero-shot mutation prediction is vital for low-resource protein engineering, yet existing protein language models (PLMs) often yield statistically confident results that ignore fundamental biophysical constraints. Currently, selecting candidates for wet-lab validation relies on manual expert auditing of PLM outputs, a process that is inefficient, subjective, and highly dependent on domain expertise. To address this, we propose Rank-and-Reason (VenusRAR), a two-stage agentic framework to automate this workflow and maximize expected wet-lab fitness. In the Rank-Stage, a Computational Expert and Virtual Biologist aggregate a context-aware multi-modal ensemble, establishing a new Spearman correlation record of 0.551 (vs. 0.518) on ProteinGym. In the Reason-Stage, an agentic Expert Panel employs chain-of-thought reasoning to audit candidates against geometric and structural constraints, improving the Top-5 Hit Rate by up to 367% on ProteinGym-DMS99. The wet-lab validation on Cas12i3 nuclease further confirms the framework's efficacy, achieving a 46.7% positive rate and identifying two novel mutants with 4.23-fold and 5.05-fold activity improvements. Code and datasets are released on GitHub (https://github.com/ai4protein/VenusRAR/).
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.