가설 트리 정제를 통한 범용 자율 연구
Toward Generalist Autonomous Research via Hypothesis-Tree Refinement
과학 발전은 탐색, 실험 및 추상의 반복적인 과정을 통해 이루어진다. 연구자들은 후보 방향을 테스트하고, 증거를 해석하며, 얻은 교훈을 후속 시도에 활용한다. 본 논문에서는 AI 에이전트가 이러한 루프를 장기간 동안 자율적으로 수행하는 방법을 연구한다. 우리는 Arbor라는 범용 자율 연구 프레임워크를 소개한다. Arbor는 장기적인 관리를 담당하는 조정자와, 짧은 시간 동안 작업을 수행하는 실행자, 그리고 가설, 산출물, 증거 및 추출된 통찰력을 시간 경과에 따라 연결하는 지속적인 트리 구조인 가설 트리 정제(HTR)를 결합한다. 조정자는 트리를 통해 전반적인 연구 전략을 관리하며, 실행자는 개별 가설을 독립적인 작업 환경에서 구현하고 테스트한다. 실험 결과가 반환되면 Arbor는 트리를 업데이트하고, 재사용 가능한 교훈을 전달하며, 탐색 범위를 개선하고, 검증된 개선 사항을 적용한다. 이러한 설계는 자율 연구를 일련의 개별 시도로부터 전략, 실행 및 증거가 시간 경과에 따라 누적되는 축적적인 프로세스로 전환시킨다. 우리는 모델 훈련, 하드웨어 엔지니어링 및 데이터 합성 등 여섯 가지 실제 연구 과제에서 Arbor를 평가했다. 그 결과, Arbor는 모든 여섯 가지 과제에서 Codex와 Claude Code보다 평균적으로 2.5배 이상의 상대적인 성능 향상을 보였다. MLE-Bench Lite에서는 GPT-5.5와 함께 Arbor가 86.36%의 Any Medal을 달성하여, 비교 대상 시스템 중 가장 뛰어난 결과를 나타냈다.
Scientific progress depends on a repeated loop of exploration, experimentation, and abstraction. Researchers test candidate directions, interpret the evidence, and carry the resulting lessons into later attempts. We study how an AI agent can run this loop autonomously over long horizons. We introduce Arbor, a general framework for autonomous research that combines a long-lived coordinator, short-lived executors, and Hypothesis Tree Refinement (HTR), a persistent tree that links hypotheses, artifacts, evidence, and distilled insights across time. The coordinator manages global research strategy over the tree, while executors implement and test individual hypotheses in isolated worktrees. As results return, Arbor updates the tree, propagates reusable lessons, refines the search frontier, and admits verified improvements. This design turns autonomous research from a sequence of local attempts into a cumulative process in which strategy, execution, and evidence are carried across time. We evaluate Arbor under Autonomous Optimization (AO), an operational setting where an agent improves an initial research artifact through iterative experimentation without step-level human supervision. Across six real research tasks in model training, harness engineering, and data synthesis, Arbor achieves the best held-out result on all six tasks, attaining more than 2.5x the average relative held-out gain of Codex and Claude Code under the same task interface and resource budget. On MLE-Bench Lite, Arbor reaches 86.36% Any Medal with GPT-5.5, the strongest result in our comparison.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.