AutoScientists: 장기간 과학 실험을 위한 자기 조직화 에이전트 팀
AutoScientists: Self-Organizing Agent Teams for Long-Running Scientific Experimentation
과학 연구는 가설 생성, 실험 설계, 실행 및 수정의 반복적인 과정을 거칩니다. 인공지능 에이전트는 이 과정의 일부를 자동화할 수 있지만, 기존 접근 방식은 일반적으로 단일 연구 경로를 따르거나 고정된 목표를 가진 중앙 계획자를 통해 협력합니다. 그 결과, 이러한 시스템은 병렬 탐색을 유지하거나 실험적 증거가 변경됨에 따라 적응하거나 장기간 실험에서 실패한 방향에 대한 지식을 보존하는 데 어려움을 겪습니다. 본 논문에서는 장기간의 계산 과학 실험을 위한 분산형 인공지능 에이전트 팀인 AutoScientists를 소개합니다. AutoScientists의 에이전트는 공유된 실험 상태를 해석하고, 유망한 가설을 중심으로 팀을 형성하며, 실험 컴퓨팅 리소스를 사용하기 전에 제안서를 비판적으로 검토하고, 성공 및 실패 사례를 공유하여 중복 탐색을 줄입니다. 동일한 실험 예산을 기준으로 AutoScientists는 생물 의학 머신 러닝, 언어 모델 훈련 최적화 및 단백질 적합성 예측 분야에서 기존 인공지능 에이전트보다 우수한 성능을 보였습니다. BioML-Bench 데이터셋(생물의료 이미징, 단백질 공학, 단일 세포 오믹스 및 신약 개발 포함)의 24개 작업에서 AutoScientists는 평균 리더보드 백분율 74.4%를 달성하여 가장 강력한 인공지능 에이전트보다 +8.33% 향상되었습니다. GPT 훈련 최적화에서는 AutoScientists가 Autoresearch보다 목표 검증 비트-바이트 비율을 1.9배 빠르게 달성했으며, 단일 에이전트 접근 방식으로는 발견할 수 없는 개선 사항을 계속 발견했습니다 (7건의 개선 사항 vs. 0건). ProteinGym 적합성 예측에서는 AutoScientists가 ACE2-Spike 결합에 대한 방법을 발견하여 현재 최고 성능 모델보다 Spearman 상관관계에서 +12.5% 향상된 결과를 얻었습니다. 수정 없이 모든 217개의 ProteinGym 실험에 적용했을 때, 동일한 방법은 이전 최고 성능 모델보다 Spearman 상관관계에서 +6.5% 향상된 결과를 보였습니다.
Scientific research proceeds through iterative cycles of hypothesis generation, experiment design, execution, and revision. AI agents can automate parts of this process, but existing approaches typically follow a single research trajectory or coordinate through a central planner with fixed objectives. As a result, they struggle to sustain parallel exploration, adapt as experimental evidence changes, or preserve knowledge of failed directions over long-running experiments. We introduce AutoScientists, a decentralized team of AI agents for long-running computational scientific experimentation. Agents interpret a shared experimental state, self-organize into teams around promising hypotheses, critique proposals before using experimental compute, and share successes and failures to reduce redundant exploration. Under matched experimental budgets, AutoScientists improves over prior AI agents across biomedical machine learning, language-model training optimization, and protein fitness prediction. On BioML-Bench, spanning biomedical imaging, protein engineering, single-cell omics, and drug discovery, AutoScientists achieves a mean leaderboard percentile of 74.4% across 24 tasks, improving over the strongest AI agent by +8.33%. On GPT training optimization, AutoScientists reaches a target validation bits-per-byte 1.9x faster than Autoresearch and continues discovering improvements from a starting champion where the single-agent approach finds none (7 vs. 0 accepted improvements). On ProteinGym fitness prediction, AutoScientists discovers a method for ACE2-Spike binding that improves over the current state-of-the-art model by +12.5% in Spearman correlation. Applied without modification across all 217 ProteinGym assays, the same method improves over the prior state of the art by +6.5% (Spearman correlation).
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.