2607.28568v1 Jul 30, 2026 cs.CL

Frontis-MA1: 머신러닝 엔지니어링 분야에서 재귀적 자기 개선을 위한 AI4AI 모델 학습

Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering

Yuchen Fan
Yuchen Fan
Citations: 1,075
h-index: 8
Kaiyan Zhang
Kaiyan Zhang
Citations: 2,223
h-index: 22
Yuxin Zuo
Yuxin Zuo
Citations: 1,188
h-index: 10
Yu Fu
Yu Fu
Citations: 223
h-index: 5
Tian-Yuan Luo
Tian-Yuan Luo
Citations: 464
h-index: 4
Dianqiao Lei
Dianqiao Lei
Citations: 4
h-index: 1
Junlin Yang
Junlin Yang
Citations: 245
h-index: 5
Che Jiang
Che Jiang
Citations: 479
h-index: 10
Can Ren
Can Ren
Citations: 47
h-index: 3
Weizhi Wang
Weizhi Wang
Citations: 0
h-index: 0
Kaikai Zhao
Kaikai Zhao
Citations: 38
h-index: 4
Hongyi Liu
Hongyi Liu
Citations: 0
h-index: 0
Kai Tian
Kai Tian
Citations: 293
h-index: 7
Zhenzhao Yuan
Zhenzhao Yuan
Citations: 147
h-index: 2
Xiaojian Lin
Xiaojian Lin
Citations: 0
h-index: 0
Li Sheng
Li Sheng
Citations: 235
h-index: 4
Rushi Qiang
Rushi Qiang
Citations: 174
h-index: 6
Guoli Jia
Guoli Jia
Citations: 287
h-index: 7
Xingtai Lv
Xingtai Lv
Citations: 956
h-index: 11
Ermo Hua
Ermo Hua
Citations: 669
h-index: 11
Youbang Sun
Youbang Sun
Citations: 779
h-index: 9
Ning Ding
Ning Ding
Citations: 97
h-index: 4
Bowen Zhou
Bowen Zhou
Citations: 956
h-index: 9
Yuru Wang
Yuru Wang
Citations: 155
h-index: 2

재귀적 자기 개선(Recursive Self-Improvement, RSI)은 인공지능 시스템이 인공지능 개발 프로세스 자체를 개선하는 것을 요구하며, 머신러닝 엔지니어링(MLE)은 이러한 능력을 연구하기 위한 구체적인 실행 가능한 테스트 환경을 제공합니다. 우리는 MLE 분야의 RSI 연구를 위한 완전한 스택의 오픈 소스 시스템인 OpenMLE를 소개합니다. 여기에는 검증 가능한 작업 환경과 실행 피드백(OpenMLE-Gym), 연산자 학습(OpenMLE-RL) 및 장기 검색(OpenMLE-Evo)이 포함됩니다. 우리는 이 플랫폼 위에서 Frontis-MA1 (35B)을 메타 진화 에이전트로 추가 학습시켜, 사후 훈련과 추론을 네 가지 기본 프로그램 진화 연산자(Draft, Improve, Debug, Crossover)를 중심으로 진행합니다. 동일한 연산자는 실행 기반의 지도 학습 및 강화 학습을 통해 모든 평가 벤치마크에 대한 데이터 중복 제거 후 학습되며, 장기 검색으로 구성되어 학습과 진화를 단일 루프에서 결합합니다. 하나의 RTX 4090 (12GB VRAM 제한)에서 각 작업에 12시간의 예산을 사용하여 MLE-Bench Lite에서 Frontis-MA1 (35B)은 OpenMLE-Evo를 통해 기본 모델의 평균 점수를 39.39%에서 60.61%로 향상시키고, OpenMLE-Evo-Max (벤치마크 독립적인 사전 경험 및 비동기 검색 사용)를 사용하여 71.21%에 도달하여 GPT-5.5 + Codex를 능가하고 GPT-5.6 Sol 및 2.8T Kimi K3에 근접하는 성능을 보입니다. 별도의 NatureBench Lite 데이터셋에서, 프레임워크와 모델 모두의 전이성이 확인되었습니다. 프레임워크를 고정하고 학습된 모델을 적용하면 SOTA 매칭 비율이 50%에서 70%로 향상되고, 모델을 고정하고 OpenMLE-Evo를 적용하면 해당 비율이 20%에서 50%로 향상됩니다. 우리는 모델 가중치와 전체 OpenMLE 스택을 공개하여 실행 가능한 AI4AI 연구를 통한 RSI에 대한 재현 가능한 연구를 지원합니다. 코드: https://github.com/FrontisAI/OpenRSI

Original Abstract

Recursive self-improvement (RSI) requires AI systems that improve the process of building AI (i.e., AI4AI); machine learning engineering (MLE) offers a concrete, executable testbed for studying this capability. We introduce OpenMLE, an open full-stack system for RSI research in MLE, spanning verifiable task environments with execution feedback (OpenMLE-Gym), operator learning (OpenMLE-RL), and long-horizon search (OpenMLE-Evo). On this stack we post-train Frontis-MA1 (35B) as a meta-evolution agent for MLE, aligning post-training and inference around four atomic program-evolution operators (Draft, Improve, Debug, Crossover): the same operators are trained via execution-grounded SFT and RL on data deduplicated against all evaluation benchmarks, then composed into long-horizon search, coupling learning and evolution in a single loop. On MLE-Bench Lite under a 12-hour per-task budget on one RTX 4090 capped at 12 GB VRAM, Frontis-MA1 (35B) improves Medal Average from 39.39% to 60.61% over its base model with OpenMLE-Evo, and reaches 71.21% with OpenMLE-Evo-Max (benchmark-independent experience priors and asynchronous search), exceeding GPT-5.5 + Codex and approaching GPT-5.6 Sol and the 2.8T Kimi K3. On held-out NatureBench Lite, both components transfer: with the framework fixed, swapping in the trained model raises Match-SOTA from 50% to 70%; with the model fixed, swapping in OpenMLE-Evo raises it from 20% to 50%. We release the model weights and the full OpenMLE stack to enable reproducible research on executable AI4AI toward RSI. Code: https://github.com/FrontisAI/OpenRSI

2 Citations
0 Influential
0 Altmetric
11.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!