TxBench-PP: 소분자 전임상 약리학 분야에서 인공지능 에이전트 성능 분석
TxBench-PP: Analyzing AI Agent Performance on Small-Molecule Preclinical Pharmacology
인공지능(AI) 에이전트는 해석 및 의사 결정 과정을 단축하여 신약 개발을 가속화할 잠재력을 가지고 있지만, 실제 적용을 위해서는 현실적인 의사 결정에 대한 신뢰성 있는 평가가 필요합니다. 본 연구에서는 소분자 전임상 약리학 분야를 위한 검증 가능한 벤치마크인 TherapeuticsBench Preclinical Pharmacology (TxBench-PP)를 소개하며, 이는 더 광범위한 TherapeuticsBench 프로젝트의 일환으로 신약 개발 단계 및 치료 모달리티 전반에 걸쳐 적용될 예정입니다. TxBench-PP는 에이전트가 문헌에서 암기된 사실이 아닌 실제 실험 데이터로부터 정확한 결론을 도출할 수 있는지 테스트합니다. 이 벤치마크는 프로그램 단계, 분석 유형 및 작업 구조를 기준으로 100개의 평가 항목으로 구성되며, 작용 기전(MoA) 및 약동학(PD) 추론, 화합물-표적 상호작용, 인과 관계 표적 검증, 개발 가능성 및 안전성, 그리고 번역 가능 효능을 포함합니다. 에이전트는 실제 워크플로우 스냅샷을 받고 코딩 환경에서 파일을 검토한 후 구조화된 답변을 제공하며, 이 답변은 결정적으로 평가됩니다. 11개의 모델과 4,800개의 시뮬레이션을 조합한 16가지 모델-하네스 구성에서, 어떤 시스템도 전임상 약리학 결정을 안정적으로 도출하지 못했습니다. 가장 강력한 구성인 Claude Opus 4.8 / Pi는 최종 목표에 대한 시도 중 59.3% (178/300; 95% CI, 51.1-67.6)를 달성했으며, 다음으로 GPT-5.5 / Pi가 55.3% (166/300; 47.0-63.6)를 기록했습니다.
Artificial intelligence (AI) agents promise to accelerate drug discovery by compressing interpretation and decision-making loops, but practical deployment requires trusted evaluation on realistic program decisions. We introduce TherapeuticsBench Preclinical Pharmacology (TxBench-PP), a verifiable benchmark for small-molecule preclinical pharmacology and the first focused slice of a broader TherapeuticsBench effort across drug-discovery stages and therapeutic modalities. TxBench-PP tests whether agents can recover accurate conclusions from real-world assay data rather than memorized facts from literature. The benchmark contains 100 evaluations indexed by program stage, assay type, and task structure, spanning mechanism-of-action (MoA) and pharmacodynamic (PD) reasoning, compound-target engagement, causal target validation, developability and safety, and translational efficacy. Agents receive realistic workflow snapshots, inspect files in a coding environment, and return structured answers graded deterministically. Across 16 model-harness configurations, comprising 11 models and 4,800 trajectories, no system reliably recovered preclinical pharmacology decisions. The strongest configuration, Claude Opus 4.8 / Pi, passed 59.3\% of endpoint attempts (178/300; 95\% CI, 51.1-67.6), followed by GPT-5.5 / Pi at 55.3\% (166/300; 47.0-63.6).
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.