AI 에이전트가 개방형 AI 연구를 수행할 수 있는가? 두 가지 사례 연구를 통한 초기 증거
Can AI agents conduct open-ended AI research? Early evidence from two case studies
폭발적인 AI 발전은 AI 에이전트가 AI 연구 자체를 자동화하는 데 달려 있습니다. 하지만, 에이전트가 개방형 AI 연구를 수행할 수 있는지에 대한 증거는 부족합니다. 현재의 평가는 일반적으로 좁고 검증 가능한 작업에 대해 에이전트를 테스트하거나, AI가 생성한 논문을 익명 심사 과정에 투입하는 방식으로 진행되는데, 이는 개방형 연구를 포함하지 않거나, 과도하고 확률적이며, 심사 품질이 좋지 않은 문제가 있습니다. 우리는 AI 연구 개발 자동화의 진척 상황을 측정하는 세 번째 방법을 제시합니다. 에이전트는 고품질의 출판되지 않은 논문의 핵심적인, 개방형 연구 질문을 맡고, 해당 논문의 원 저자들이 에이전트의 결과물을 평가합니다. 우리는 이를 '섀도우 평가(shadow evaluation)'라고 부릅니다. 우리는 최첨단 에이전트를 사용하여 NeurIPS 2026에 제출될 예정인 두 편의 출판되지 않은 논문에 대해 섀도우 평가를 수행했으며, 에이전트에 6일 동안 수천 달러 상당의 컴퓨팅 자원을 제공했습니다. 에이전트는 인간의 도움 없이 모든 엔지니어링 작업을 완료했지만, 연구 질문에 대한 실질적인 진전을 이루지는 못했습니다. 그 결과, 두 논문 모두 원 저자에 의해 명확하게 거부되었습니다. 우리는 다음과 같은 다섯 가지 반복되는 실패 요인을 확인했습니다: 출판 가능한 연구 수준에 대한 판단 오류, 연구 설계의 부족한 부분에 대한 비창의적인 대응, 막다른 길에서 벗어나지 못하는 문제, 자원 활용 능력 부족, 그리고 지시 사항의 왜곡. 두 번째 모델과 스캐폴드를 사용한 견고성 검증에서도 이러한 실패가 재현되었습니다. 우리는 전문가 심사 결과, 설문 조사 응답, 에이전트 저장소 및 로그를 공개합니다. 우리의 연구 결과는 현재의 에이전트가 AI 연구의 엔지니어링 부분은 수행할 수 있지만, 연구 생명 주기의 중요한 부분에서는 어려움을 겪는다는 초기 증거를 제공합니다.
Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carry out open-ended AI research is thin. Current evaluations either test agents on narrow, verifiable tasks, which excludes open-ended research, or submit AI-generated papers to blind peer review, which is overstretched, stochastic, and suffers from poor review quality. We introduce a third way to measure progress towards AI R\&D automation. An agent takes on the central, open-ended research question of a high-quality unpublished paper, and the paper's original authors grade its output. We call these shadow evaluations. We ran shadow evaluations on two unpublished NeurIPS 2026 submissions, giving frontier agents six days and thousands of dollars of compute. The agents completed all of the engineering without human help, yet could not make substantial progress towards answering the research questions. As a result, both papers were unambiguously rejected by the authors. We identify five recurring failure modes: poor judgment about the bar for publishable research, uncreative responses to shortcomings in the research design, ineffective backtracking from dead ends, poor resource awareness, and instruction drift. A robustness check with a second model and scaffold reproduced these failures. We release the expert reviews, survey responses, agent repositories, and logs. Our results provide early evidence that today's agents can do the engineering of AI research, but struggle with critical parts of the research lifecycle.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.