2607.26587v1 Jul 29, 2026 cs.MA

단 한 번의 실행은 하나의 아이디어가 아니다: 자동화된 연구에서의 구현 추첨 (Implementation Lottery)

One Run Is Not an Idea: The Implementation Lottery in Automated Research

Ji Zeng
Ji Zeng
Citations: 35
h-index: 4
Xiaochuan Li
Xiaochuan Li
Citations: 1,008
h-index: 5
Jingjie Ning
Jingjie Ning
Citations: 55
h-index: 2
Chenyan Xiong
Chenyan Xiong
Citations: 74
h-index: 3
Shan Zhong
Shan Zhong
Citations: 415
h-index: 9

자동화된 연구 시스템은 실험 결과를 사용하여 결과물을 생성하고, 어떤 아이디어를 유지하고, 이전하며, 발전시킬지 결정합니다. 그러나 단 한 번의 실행은 특정 아이디어의 하나의 구현을 평가하는 것입니다. 이러한 구현 수준의 점수를 부모 메커니즘에 대한 증거로 사용하는 것은 "구현 추첨(implementation lottery)"이라는 현상을 야기하며, 이 경우 아이디어 수준의 결론은 어떤 가능한 구현이 선택되었는지에 따라 달라집니다. 하나의 실행 결과가 메커니즘에 대한 믿음을 업데이트할 때마다 이러한 불일치는 구조적으로 발생합니다. 우리는 이러한 불일치의 크기를 추정했습니다. "아이디어 신뢰도 감사(Idea Reliability Audit)"는 후보 카드들을 검증하고 고정하며, 새로운 세션의 구현을 샘플링하고, 결과에 영향을 받지 않는 품질 레이블을 사용하고, 저장된 결과물을 재실행하여 "아이디어 신뢰도"를 측정합니다. 이 방법은 아이디어의 ICC (분산 중복성)와 leave-one-implementation-out (LOO) 방식을 사용하여 승자 결정의 일관성을 평가합니다. 기존 연구는 일반적으로 동일한 작업을 반복하지만, 우리는 아이디어를 반복했습니다. 13개의 표 형식 작업과 두 가지 코딩 에이전트 환경에서 총 312개의 실험을 수행한 결과, 구현 변동은 동일한 결과물에 대한 재실행 변동보다 각각 5배 이상 및 10배 이상 높았으며, 하나의 구현에서 선택된 "승자"가 다른 두 구현의 평균에서 결정된 "승자"와 25.6% 및 43.6%의 결정에서 달랐습니다. 이러한 불일치는 카드 수준의 필터링을 거쳐 결과에 영향을 받지 않는 검토 규칙 하에서도 유지됩니다. 또한, 세 가지 재료 회귀 워크플로우를 대상으로 하는 탐색적 진단 분석에서는 결정적인 평가 기준을 사용했음에도 불구하고 구현 변동이 분해 과정에서 지배적인 역할을 한다는 사실이 확인되었습니다. 이러한 결과는 아이디어 신뢰도를 최상의 N개 결과물의 유틸리티와 구별합니다. 점수가 아이디어 수준의 분기, 이전 또는 연구 기억에 영향을 미치기 전에, 증거는 여러 구현을 기반으로 해야 합니다.

Original Abstract

Automated research systems use experimental scores both to deliver artifacts and to decide which ideas to retain, transfer, and pursue. Yet one run scores one implementation of an idea. Crediting that realization-level score as evidence about the parent mechanism creates the \emph{implementation lottery}, in which an idea-level conclusion depends on which plausible implementation was sampled. The mismatch is structural whenever one run updates beliefs about a mechanism. We estimate its magnitude. The \emph{Idea Reliability Audit} measures \emph{idea reliability} by validating and freezing candidate cards, sampling fresh-session implementations, using outcome-blind fidelity labels, and rerunning saved artifacts. It reports idea ICC and leave-one-implementation-out (LOO) winner reversal. Prior work generally repeats the task; we repeat the idea. Across 312 assignments on 13 tabular tasks and two coding-agent setups, implementation variance was more than five and ten times same-artifact rerun variance, respectively, and the winner from one implementation draw differed from the winner under the other-two mean in 25.6\% and 43.6\% of decisions. Reversal survives card-level filtering under two outcome-blind review rules. An exploratory diagnostic on three materials-regression workflows with a deterministic evaluator also finds implementation variation dominating the decomposition. These findings distinguish idea reliability from best-of-$N$ artifact utility. Before a score guides idea-level branching, transfer, or research memory, evidence should cover multiple implementations.

0 Citations
0 Influential
4.5 Altmetric
22.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!