2606.30182v1 Jun 29, 2026 cs.AI

MirrorCode: 인공지능이 프로그램의 동작만으로 전체 프로그램을 재구성할 수 있습니다.

MirrorCode: AI can rebuild entire programs from behavior alone

F. Brand
F. Brand
Citations: 74
h-index: 4
D. O’Connell
D. O’Connell
Citations: 52
h-index: 4
Tomasz Adamczewski
Tomasz Adamczewski
Citations: 59
h-index: 5
David Owen
David Owen
Citations: 182
h-index: 3
David Rein
David Rein
METR (Model Evaluation and Threat Research)
Citations: 3,073
h-index: 6
Giles Edkins
Giles Edkins
Citations: 69
h-index: 2
Allen G. Hart
Allen G. Hart
Citations: 290
h-index: 6

인공지능 모델은 자율 코딩 능력에서 빠르게 발전하고 있으며, 이는 벤치마크 성과 및 AI가 C 컴파일러를 구현하는 등 일회성 시연을 통해 입증됩니다. 그러나 기존의 코딩 벤치마크는 주로 짧은 작업에 초점을 맞추고 있으며, 일회성 시연은 종종 인간의 지침을 포함하기 때문에 체계적인 비교가 어렵고, 모델 간 표준화 또는 반복도 이루어지지 않습니다. 이러한 문제를 해결하기 위해, 저희는 전체 소프트웨어 프로젝트를 재구현하는 데 기반한 장기 코딩 벤치마크인 MirrorCode를 소개합니다. MirrorCode에서 AI 에이전트는 기존 프로그램의 소스 코드에 접근할 수 없는 상태에서 해당 프로그램의 기능을 복제해야 합니다. AI 솔루션은 원본 프로그램과 정확히 동일한 결과를 내야 하며, 여기에는 테스트 데이터셋 전체를 포함한 숨겨진 테스트도 포함됩니다. MirrorCode는 Unix 유틸리티, 데이터 직렬화 및 쿼리 도구, 생물정보학, 인터프리터, 정적 분석, 암호학 및 압축 등 다양한 분야의 25개의 대상 프로그램을 포함합니다. 기존의 AI 모델은 이미 복잡한 소프트웨어를 재구현할 수 있으며, 가장 뛰어난 성능을 보이는 모델은 이 벤치마크에서 56%의 정확도를 기록했습니다. 예를 들어, AI는 16,000줄로 구성된 생물정보학 도구인 gotree를 재구현할 수 있는데, 이는 저희가 생각하기에 인간 엔지니어가 몇 주 동안 걸릴 작업입니다. 그러나 성능의 한계를 연구하려면 일반적인 벤치마크보다 더 큰 연산 자원이 필요하며, 예를 들어 대규모 작업에 대한 단일 시도에는 약 2,600달러의 비용이 들고 19일이 소요됩니다. 저희는 AI 에이전트가 이미 장기적인 소프트웨어 엔지니어링 작업을 수행할 수 있으며, 특히 요구 사항이 정확하게 명시된 경우 더욱 그렇다는 것을 보여줍니다. 더 넓은 관점에서 볼 때, 저희의 연구 결과는 자율 에이전트가 지속적으로 발전함에 따라 인공지능이 소프트웨어 엔지니어링 분야에 혁신적인 영향을 미칠 것이라는 점을 시사합니다.

Original Abstract

AI models are rapidly improving at autonomous coding, as shown by benchmark progress and one-off demonstrations such as AI implementing a C compiler. However, existing coding benchmarks tend to focus on shorter tasks, and one-off demonstrations are hard to compare systematically because they often have some human guidance, and are not standardized or repeated across models. To address these challenges, we introduce MirrorCode, a long-horizon coding benchmark based on reimplementing entire software projects. In MirrorCode, AI agents must replicate the functionalities of an existing program, without access to its source code. AI solutions must match the original program's output exactly on end-to-end tests, including held-out tests. MirrorCode's 25 target programs span different areas of computing: Unix utilities, data serialization and query tools, bioinformatics, interpreters, static analysis, cryptography, and compression. Existing AI models can already reimplement complex software, with the strongest model scoring 56% across the benchmark. For example, AI can reimplement gotree, a 16,000-line bioinformatics toolkit - a task that we believe would take weeks for a human engineer. However, studying the frontier of performance requires a larger inference budget than typical benchmarks, for example, \$2,600 over 19 days for a single attempt on a large task. We show that AI agents can already complete long-horizon software engineering tasks, especially when requirements are precisely specified. More broadly, our work suggests AI will have transformative effects on software engineering, as autonomous agents continue to improve.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!