솔루션 중심 검색을 넘어: 자율 머신러닝 엔지니어링을 위한 적응적 탐색 및 지식 수정
Beyond Solution-Centric Search: Adaptive Inquiry and Knowledge Revision for Autonomous ML Engineering
머신러닝 엔지니어링과 같은 장기적인 자율 연구 과제는 시스템이 제한된 예산 하에서 상호 의존적인 결정을 내리도록 요구합니다. 기존의 LLM 기반 에이전트는 일반적으로 트리, 그래프 또는 체인 구조를 통해 후보 솔루션을 개선하는데, 이는 검색 프로세스가 정보 획득 및 관리 방식을 결정한다는 의미입니다. 우리는 이러한 설계를 '솔루션 중심 검색'이라고 부르며, 대신 '정보 패러다임'을 제안합니다. 이 패러다임에서는 시스템의 과제 이해도를 나타내는 진화하는 정보 상태가 솔루션 개선을 안내합니다. 우리는 이 패러다임을 탐색-수정 루프인 Iris에 구현했습니다. 정보 획득을 위해 Iris는 현재 정보 상태에서 로컬 액션 플랜을 생성하고, 유지된 솔루션을 수정하지 않고 의사 결정에 중요한 미지의 영역을 탐색하기 위한 인지적 행동을 사용합니다. 정보 관리를 위해 Iris는 실험 결과를 과제 지식으로 종합하며, 이 지식은 명시적인 범위와 상태를 가진 수정 가능한 주장을 포함합니다. 시스템은 새로운 증거가 도착하면 이러한 지식을 업데이트하고, 각 의사 결정 맥락을 원시 증거, 구조화된 요약 또는 필요한 수준의 세부 정보로 과제 지식에서 구성합니다. Iris는 MLE-Bench에서 12시간의 예산 하에 64.9%의 상위 수상률을 달성했으며, 이는 비교 대상 시스템 중 가장 높은 수치입니다. 또한 Iris는 하네스 엔지니어링 및 모델 후처리 등 네 가지 과제에서 다양한 영역으로 일반화되는 능력을 보여줍니다.
Long-horizon autonomous research tasks such as machine learning engineering require systems to make interdependent decisions under a limited budget. Existing LLM-based agents typically organize candidate-solution improvement through tree, graph, or chain structures, meaning that the search process determines how information is acquired and managed. We call this design solution-centric search and propose instead the information paradigm, in which an evolving information state represents the system's understanding of the task and guides solution improvement. We instantiate this paradigm in Iris, an inquiry-revision loop. For information acquisition, Iris generates local action plans from the current information state and uses epistemic actions to probe decision-critical unknowns without modifying the retained solution. For information management, Iris synthesizes observations across experiments into task knowledge composed of revisable claims with explicit scope and status. It updates this knowledge as new evidence arrives and constructs each decision context from raw evidence, structured summaries, or task knowledge at the required level of detail. On MLE-Bench, Iris attains a 64.9% any-medal rate under a 12-hour budget, the highest among compared systems. Across four tasks spanning harness engineering and model post-training, Iris also demonstrates cross-domain generalization.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.