2607.18235v1 Jul 20, 2026 cs.CL

자동화된 탐색 방법론은 보편적으로 우수한 솔루션이 존재하지 않는다

Automated Discovery Has No Universally Superior Harness

Leshem Choshen
Leshem Choshen
Citations: 250
h-index: 9
Akshat Gupta
Akshat Gupta
Citations: 53
h-index: 4
G. Anumanchipalli
G. Anumanchipalli
Citations: 4,010
h-index: 27
Alexander Lu
Alexander Lu
Citations: 6
h-index: 2
Jermaine Lei
Jermaine Lei
Citations: 0
h-index: 0

OpenEvolve 및 TTT-Discover와 같은 자율 탐색 시스템은 종종 범용적인 도구로 사용됩니다. 그러나 실제로 이러한 시스템들은 아카이브, 부모 선택, 탐색 전략, 예산 할당 등 다양한 설계 요소를 하나의 통합된 방식으로 결합한 복합 시스템입니다. 탐색 과정은 비용이 많이 들고 본질적으로 확률적이기 때문에, 기존의 탐색 도구들은 종종 충분히 독립적인 실험을 통해 주요 방법론적 개선 사항과 실행 간의 변동성을 구별하지 못합니다. 본 연구에서는 OpenEvolve 스타일의 진화 탐색 및 TTT-Discover 탐색 도구를 구성 요소로 분해하고, 30개의 동일한 예산을 사용하는 다양한 탐색 도구를 12쌍의 모델 문제에 대해 310만 번 이상의 LLM 실행과 반복적인 통계 분석을 통해 체계적으로 평가했습니다. 연구 결과는 탐색 도구가 일반화 성능 문제를 가지고 있음을 보여줍니다. 즉, 평가된 모델-문제 쌍 전반에 걸쳐 특정 탐색 도구가 일관되게 우수한 성능을 보이지 않으며, OpenEvolve의 변형은 일반적으로 더 간단한 대안보다 성능이 낮습니다. 따라서, 탐색 도구 선택은 '보편적인 솔루션'이라기보다는 하이퍼파라미터로 간주되어야 하며, 특정 문제와 기본 모델에 맞춰 조정되어야 합니다. 또한, 초기 탐색 진행 상황이 최종 성능을 예측한다는 것을 발견했으며, 이를 바탕으로 여러 탐색 도구를 동시에 시작하고, 성과가 낮은 부분 실행 결과를 제거하며, 더 나은 결과를 보이는 도구에 컴퓨팅 자원을 재할당하는 적응적 할당 실험을 수행했습니다. 이 실험은 임의로 선택된 고정된 탐색 도구를 사용하는 방법이나 비적응형 탐색 도구 조합보다 우수한 성능을 보였습니다. 이러한 결과는 고정된 탐색 도구 선택에서 초기 성능에 기반한 온라인 적응 방식으로 전환하는 것을 시사합니다. 또한, 모든 실행 결과를 제공하며, 각 모델-문제 쌍에 대한 기본 기준 분포를 포함하여 향후 탐색 도구 제안을 위한 재사용 가능한 통계 인프라로 활용할 수 있도록 합니다.

Original Abstract

Autonomous discovery systems such as OpenEvolve and TTT-Discover are often used as general-purpose harnesses. However, in practice these are composite systems combining several design choices about archives, parent selection, exploration, and budget allocation into a single recipe. Because discovery runs are expensive and inherently stochastic, existing harnesses are often compared using too few independent trials to distinguish key methodological improvements from run-to-run variance. We systematically decompose OpenEvolve-style evolutionary search and the TTT-Discover search harness into its constituent components and systematically evaluate 30 budget-matched harnesses across 12 model-problem pairs using more than 3.1 million LLM rollouts and repeated-trial statistical analysis. Our results show that discovery harnesses have a generalization problem: No fixed harness is reliably superior across the evaluated model-problem pairs, and variants of OpenEvolve generally underperform simpler alternatives. Thus, harness choice is better viewed as a hyperparameter rather than as a universal recipe, and should be tailored to the specific problem and underlying model. We also find that early discovery progress predicts final performance, and use this property to present a budget-matched adaptive-allocation experiment that starts multiple harnesses, prunes weak partial runs, and reallocates compute to stronger survivors, outperforming both commitment to a randomly sampled fixed harness and a non-adaptive harness ensemble. Together, these results motivate shifting from fixed harness selection to online adaptation guided by early performance. We release all run pools including baseline null distributions for every model-problem pair as reusable statistical infrastructure against for future harness proposals.

0 Citations
0 Influential
13.5 Altmetric
67.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!