2606.09663v1 Jun 08, 2026 cs.AI

0에서 1로, 그리고 1에서 N으로: 메타 AI의 재귀적 자기 설계에 대한 재현 가능한 공학적 증거

From 0-to-1 to 1-to-N: Reproducible Engineering Evidence for MetaAI Recursive Self-Design

Hongzhi Li
Hongzhi Li
Citations: 4
h-index: 1
Dun Li
Dun Li
Citations: 28
h-index: 4
Jiatao Li
Jiatao Li
Peking University
Citations: 24
h-index: 4

재귀적 자기 설계는 인공지능 시스템이 구축, 평가 및 개선되는 방식 자체를 수정하는 인공지능 지원 개량 방식을 의미합니다. 본 논문에서는 MetaAI를 성숙한 패러다임으로 간주하지 않고, 인간의 초기 설계에 기반하여 인공지능이 확장하는 개발 패턴을 지칭하는 용어로 사용합니다. 우리는 검증 가능한 대상 시스템, 메타 수준 수정기, 피드백 기반 선택 및 재귀적 연속이라는 네 가지 기준을 가진 운영 증거 프레임워크를 제안합니다. 그런 다음 Darwin Goedel Machine (DGM), STOP, Goedel Agent 및 ShinkaEvolve와 같은 공개 시스템을 이러한 기준에 따라 분석합니다. DGM은 현재 보고된 가장 직접적인 증거를 제공합니다. 발표된 결과는 80번의 반복 후 SWE-bench Verified에서 20%에서 50%로, full Polyglot에서 14.2%에서 30.7%로 성능이 향상되었으며, 분석 결과 개방형 탐색과 자기 개선 모두에 기여하는 것으로 나타났습니다. 마지막으로, 재현 가능한 HumanEval 기반 프로토콜 및 코드를 제공하는 MetaAI-Mini를 제시합니다. 이 빌드에는 완료된 모델 실행이 포함되어 있지 않으므로 MetaAI-Mini는 실험 결과가 아닌 프로토콜로 보고됩니다.

Original Abstract

Recursive self-design refers to AI-assisted modification of the mechanisms by which an AI system is built, evaluated, and improved. This paper treats MetaAI not as a mature paradigm, but as a working term for a human-seeded, AI-expanded development pattern in which the design space itself becomes a target of modification. We propose an operational evidence framework with four criteria: inspectable target system, meta-level modifier, feedback-directed selection, and recursive continuation. We then map public systems, including Darwin Goedel Machine (DGM), STOP, Goedel Agent, and ShinkaEvolve, against these criteria. DGM provides the most direct currently reported evidence: its published results show improvement from 20% to 50% on SWE-bench Verified and from 14.2% to 30.7% on full Polyglot after 80 iterations, with ablations suggesting that both open-ended exploration and self-improvement contribute. Finally, we provide MetaAI-Mini, a reproducible HumanEval-based protocol and codebase. Because no completed model run is included in this build, MetaAI-Mini is reported as a protocol rather than as an experimental result.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!