임상 시험 전략 학습: 의사 결정 에이전트를 위한 오프라인 정책 학습
Learning Clinical-Trial Strategy: Offline Policy Training for Decision Agents
임상 개발은 불확실성 하에서 순차적인 의사 결정을 요구하며, 연구 제약 회사는 다양한 증거를 바탕으로 실험 포트폴리오를 계획해야 합니다. 본 연구에서는 종양학 임상 개발을 오프라인 의사 결정 문제로 설정하여, 에이전트가 특정 시점에서 사용 가능한 정보를 기반으로 종양 치료제의 6개월간의 시험 포트폴리오를 예측하도록 합니다. 이를 위해, 임상시험 등록 정보, 규제 검토 자료, 제약 회사 제출 서류, 이용 데이터 및 역학 데이터를 포함한 31,700건의 다양한 공개 데이터를 수집하여 45개의 역사적 프로그램에 걸쳐 881개의 오프라인 의사 결정 에피소드로 구성된 시계열 데이터 세트를 구축했습니다. 본 연구에서는 행동 복제(behavioral cloning), 보상 가중 행동 복제(reward-weighted behavioral cloning), 학습 기반 보상 학습(learned-reward training) 및 가치 기반 암묵적 Q-러닝(value-based implicit Q-learning)의 네 가지 오프라인 목표를 비교했으며, 공통된 시계열 정보 검색 구조를 사용하는 최첨단 LLM 에이전트 4개를 사용하여 평가했습니다. 오프라인으로 학습된 모델은 미세 조정되지 않은 기준 모델보다 성능이 우수했으며, 특히 2025년 8월 이후의 데이터 세트를 사용한 평가에서 두드러진 결과를 보였습니다. 보상 가중 행동 복제가 가장 좋은 성능을 나타냈으며, 해당 지표에서 가장 높은 성능을 보인 도구 기반 에이전트에 비해 각각 46.2%의 indication F1 점수와 14.2%의 strict F1 점수를 기록했습니다 (기준 모델은 각각 25.0% 및 2.1%). 이러한 결과는 구조화된 오프라인 학습을 통해 에이전트가 임상 시험 계획 수립 능력을 갖추도록 할 수 있음을 시사합니다.
Clinical development is sequential decision-making under uncertainty, where a sponsor must plan a portfolio of experiments from heterogeneous evidence. We study this setting by framing oncology clinical development as an offline decision-making problem in which an agent predicts the next six-month trial portfolio of an oncology drug program from information available at the decision date. To support this, we construct a temporal dataset that combines 31.7k heterogeneous public data records, including trial registries, regulatory reviews, sponsor filings, utilization data, and epidemiology, into 881 offline decision episodes across 45 historical programs. We compare four offline objectives: behavioral cloning, reward-weighted behavioral cloning, learned-reward training, and value-based implicit Q-learning against four frontier LLM agents that share a common date-gated retrieval scaffold across held-out drug, sponsor, drug-class, and temporal splits. Models trained offline outperform the non-fine-tuned baselines, particularly in the post-August 2025 contamination-clean holdout. Reward-weighted behavioral cloning performs the best, obtaining 46.2% indication F1 and 14.2% strict F1 against 25.0% and 2.1%, respectively, for the best-performing tool agent on each metric. These results suggest that structured offline learning can teach agents to plan clinical experiments.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.