2606.19245v1 Jun 17, 2026 cs.AI

TxBench-PP: 소분자 전임상 약리학 분야에서 인공지능 에이전트 성능 분석

TxBench-PP: Analyzing AI Agent Performance on Small-Molecule Preclinical Pharmacology

Kenny Workman
Kenny Workman
Citations: 5
h-index: 2
Hannah Le
Hannah Le
Citations: 5
h-index: 2
Timothy Proctor
Timothy Proctor
Citations: 11
h-index: 2
R. Ramasamy
R. Ramasamy
Citations: 0
h-index: 0
Alex Urrutia
Alex Urrutia
Citations: 0
h-index: 0
M. Yazdani
M. Yazdani
Citations: 6
h-index: 2

인공지능(AI) 에이전트는 해석 및 의사 결정 과정을 단축하여 신약 개발을 가속화할 잠재력을 가지고 있지만, 실제 적용을 위해서는 현실적인 의사 결정에 대한 신뢰성 있는 평가가 필요합니다. 본 연구에서는 소분자 전임상 약리학 분야를 위한 검증 가능한 벤치마크인 TherapeuticsBench Preclinical Pharmacology (TxBench-PP)를 소개하며, 이는 더 광범위한 TherapeuticsBench 프로젝트의 일환으로 신약 개발 단계 및 치료 모달리티 전반에 걸쳐 적용될 예정입니다. TxBench-PP는 에이전트가 문헌에서 암기된 사실이 아닌 실제 실험 데이터로부터 정확한 결론을 도출할 수 있는지 테스트합니다. 이 벤치마크는 프로그램 단계, 분석 유형 및 작업 구조를 기준으로 100개의 평가 항목으로 구성되며, 작용 기전(MoA) 및 약동학(PD) 추론, 화합물-표적 상호작용, 인과 관계 표적 검증, 개발 가능성 및 안전성, 그리고 번역 가능 효능을 포함합니다. 에이전트는 실제 워크플로우 스냅샷을 받고 코딩 환경에서 파일을 검토한 후 구조화된 답변을 제공하며, 이 답변은 결정적으로 평가됩니다. 11개의 모델과 4,800개의 시뮬레이션을 조합한 16가지 모델-하네스 구성에서, 어떤 시스템도 전임상 약리학 결정을 안정적으로 도출하지 못했습니다. 가장 강력한 구성인 Claude Opus 4.8 / Pi는 최종 목표에 대한 시도 중 59.3% (178/300; 95% CI, 51.1-67.6)를 달성했으며, 다음으로 GPT-5.5 / Pi가 55.3% (166/300; 47.0-63.6)를 기록했습니다.

Original Abstract

Artificial intelligence (AI) agents promise to accelerate drug discovery by compressing interpretation and decision-making loops, but practical deployment requires trusted evaluation on realistic program decisions. We introduce TherapeuticsBench Preclinical Pharmacology (TxBench-PP), a verifiable benchmark for small-molecule preclinical pharmacology and the first focused slice of a broader TherapeuticsBench effort across drug-discovery stages and therapeutic modalities. TxBench-PP tests whether agents can recover accurate conclusions from real-world assay data rather than memorized facts from literature. The benchmark contains 100 evaluations indexed by program stage, assay type, and task structure, spanning mechanism-of-action (MoA) and pharmacodynamic (PD) reasoning, compound-target engagement, causal target validation, developability and safety, and translational efficacy. Agents receive realistic workflow snapshots, inspect files in a coding environment, and return structured answers graded deterministically. Across 16 model-harness configurations, comprising 11 models and 4,800 trajectories, no system reliably recovered preclinical pharmacology decisions. The strongest configuration, Claude Opus 4.8 / Pi, passed 59.3\% of endpoint attempts (178/300; 95\% CI, 51.1-67.6), followed by GPT-5.5 / Pi at 55.3\% (166/300; 47.0-63.6).

1 Citations
0 Influential
1 Altmetric
6.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!