2606.22731v1 Jun 22, 2026 cs.AI

분자 특성 예측을 위한 폐루프 자동 연구: 일반화된 개선 사항 발견 및 검증

Closed-loop Auto Research for Molecular Property Prediction: Discovering and Certifying Generalizable Improvements

Jiaqi Zeng
Jiaqi Zeng
Citations: 1,033
h-index: 12
Xiaochuan Li
Xiaochuan Li
Citations: 1,008
h-index: 5
Guolin Ke
Guolin Ke
Citations: 817
h-index: 13
Jingjie Ning
Jingjie Ning
Citations: 55
h-index: 2
Chenyan Xiong
Chenyan Xiong
Citations: 74
h-index: 3

폐루프 자동 연구는 기존의 고정 데이터셋 기반 머신러닝 방식을 확장하여, 언어 모델 에이전트가 표현 방식과 모델 코드를 수정하고 외부 증거를 확보함으로써 연구 워크플로우 자체를 변경합니다. 분자 특성 예측은 다양한 작은 목표 지점을 포함합니다. 본 연구에서는 이러한 작업 공간이 검증 데이터셋을 선택하는 것 이상의 일반화된 개선을 가져다주는지 질문합니다. 본 연구는 기능, 모델, 외부 증거의 세 가지 자동 연구 축을 정의하고, 파일 수준의 차단 기능을 통해 각 축이 기준 성능에서 얼마나 기여하는지를 분석합니다. 세 개의 벤치마크 스위트 내 총 36개의 목표 지점에 대해, 선택된 모든 구성은 검색 과정에서 라벨 정보를 전혀 활용하지 않은 별도의 테스트 데이터셋에서 한 번씩 평가됩니다. 각 목표 지점의 최적 검증 축을 활용하는 파이프라인은 평균적으로 0.013, 0.011, 0.042의 양호한 테스트 성능 향상을 보였으며, 일반화 가능한 축은 스위트에 따라 달랐습니다 (TDC 데이터셋에서는 데이터, Polaris에서는 모델, MoleculeNet에서는 기능 및 모델). 가장 큰 모델 검색 성능 향상은 검증 데이터셋에서 0.041이었지만, 테스트 데이터셋에서는 0.003으로 감소했습니다. 반면, 수집된 외부 데이터는 검증 데이터셋에서 0.022의 성능을 보였지만, 테스트 데이터셋에서는 -0.019로 나타나 일반화 가능성이 부족함을 보여줍니다. 외부 CYP2C9 기질 성능 및 약물 반감기를 향상시키는 데 사용된 수집된 데이터는 테스트 구조의 64%에서 89%에 이르는 중복 파일을 제거하는 필터를 통해 관리되었으며, 이는 일반화를 위한 필수적이지만 충분하지 않은 요소였습니다. 대조군으로 사용된 자동 머신러닝 시스템은 에이전트가 수행한 코드 수준의 모델 변경 사항을 재현하지 못했으며, 성능은 0.042에 비해 0.006에 불과했습니다. 또한, 본 연구의 파이프라인은 공유 학습 데이터셋에서 사용된 84M 파라미터의 사전 학습된 3D 모델과 경쟁력 있는 성능을 보였습니다. 본 연구는 분자 특성 예측 분야에 국한되지만, 발견 과정과 별도의 테스트를 통한 검증을 분리하는 것은 특정 도메인에 관계없이, 별도 데이터셋의 지표를 최적화하는 폐루프 시스템에 대한 일반적인 교훈을 제공합니다.

Original Abstract

Closed-loop Auto Research extends automated machine learning from fixed-dataset fitting to changing the research workflow, with language-model agents editing representations and model code and acquiring external evidence. Molecular property prediction spans many small endpoints. We ask whether this action space yields improvements generalizing beyond the validation signal selecting them. We isolate three Auto Research axes, features, models, and external evidence, under a file-level ablation lock attributing each gain to one axis over a strong baseline. Across 36 endpoints in three benchmark suites we score each selected configuration once on a held-out test whose labels the search never read. A routed pipeline taking each endpoint's best validation axis reaches positive held-out gains of 0.013, 0.011, and 0.042, the transferable axis differing by suite, data on TDC, model on Polaris, feature and model on MoleculeNet. The largest model-search gain falls from 0.041 on validation to 0.003 on test, while curated data reaches 0.022 but negative 0.019 on test, two non-transfer signatures. Curated external data raises held-out CYP2C9-substrate performance by 0.17 and half-life by 0.08, admitted through a contamination filter rejecting same-source files overlapping 64 to 89 percent of test structures, necessary but not sufficient for transfer. A matched-trial automated machine learning control did not reproduce the agent's code-level model intervention, reaching 0.006 against 0.042, and the pipeline stays competitive with an 84M-parameter pretrained 3D model on the shared training split. The experiments stay within molecular property prediction, but separating discovery from held-out certification is a domain-agnostic lesson for any closed-loop system optimising a proxy for a held-out quantity.

2 Citations
0 Influential
6.5 Altmetric
34.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!