2607.27146v1 Jul 29, 2026 cs.SE

MindForge: 소스 코드 없이 프로그램 합성 기술을 활용하여 소규모 언어 모델에게 소프트웨어 공학의 전체 라이프사이클을 교육하는 방법

MindForge: Teaching Small Language Models Whole-Life-Cycle Software Engineering via Source-Free Program Synthesis

Boyuan Chen
Boyuan Chen
Citations: 151
h-index: 7
Yihao Chen
Yihao Chen
Citations: 38
h-index: 3
Shi Chang
Shi Chang
Citations: 8
h-index: 2
Feng Lin
Feng Lin
Citations: 26
h-index: 2
Khaled Chawa
Khaled Chawa
Citations: 0
h-index: 0
Shaowei Wang
Shaowei Wang
Citations: 7
h-index: 1
Ahmed E. Hassan
Ahmed E. Hassan
Citations: 175
h-index: 7

코딩 에이전트는 기존 코드를 수정하는 소프트웨어 엔지니어링 작업에서 상당한 발전을 이루었으며, 버그 수정 및 기능 구현 등이 그 예입니다. 그러나 처음부터 완전한 프로그램을 구축하는 것은 여전히 주요 과제이며, ProgramBench에서 평가된 최첨단 모델조차도 1% 미만의 작업을 완전히 해결합니다. 이러한 어려움 중 하나는 소프트웨어 엔지니어링 라이프사이클 전체를 포괄하는 확장 가능한 학습 환경이 부족하다는 점입니다. 기존의 환경 구축 프레임워크는 일반적으로 소프트웨어 개발의 단일 단계에만 초점을 맞추기 때문입니다. 이러한 격차를 해소하기 위해, 저희는 MindForge라는 자동화된 파이프라인을 소개합니다. MindForge는 오픈 소스 명령줄 프로그램을 컴파일된 실행 파일과 문서만을 제공하는 소스 코드 없는 환경으로 변환합니다. MindForge를 사용하여 ProgramBench에 포함되지 않은 저장소에서 학습 환경을 구축하고, GLM-5.2를 교사 에이전트로 사용하는 프로그램 합성 경로로 구성된 고품질 데이터 세트를 생성했습니다. 이러한 경로를 사용하여 Qwen3.6-27B를 미세 조정하면 ProgramBench의 평균 테스트 통과율이 37.98%에서 49.51%로 향상되어, 훨씬 더 큰 최첨단 모델에 버금가는 성능을 달성합니다. 더욱이, 미세 조정된 모델은 장기적인 저장소 생성 및 번역, 버그 수정, 기능 구현, 그리고 서로 다른 언어 간의 문제 해결 등 7가지 새로운 소프트웨어 엔지니어링 벤치마크에서 기본 모델보다 일관되게 성능이 향상되었으며, RepoZero-C2Rust에서는 31.00점, DeepSWE에서는 14.16점, NL2Repo-Bench (테스트 포함/미포함)에서는 각각 10.70/4.56점, SWE-bench Verified에서는 5.04점, SWE-bench Pro에서는 5.93점, SWE-bench Multilingual에서는 5.22점, 그리고 FeatBench에서는 4.94점의 절대적인 성능 향상을 보였습니다.

Original Abstract

Coding agents have made substantial progress on software engineering tasks that modify existing codebases, including bug fixing and feature implementation. However, constructing a complete program from scratch remains a major challenge: even the frontier models evaluated on ProgramBench fully resolve fewer than 1% of tasks. One obstacle is the lack of scalable training environments for this from-scratch setting, spanning the whole software engineering life cycle, as existing environment-construction frameworks focus only on a single phase in software development. To address this gap, we introduce MindForge, an automated pipeline that converts open-source command-line programs into source-free environments that expose only a compiled reference executable and its documentation. Using MindForge, we construct training environments from repositories disjoint from those in ProgramBench, and curate a high-quality data recipe consisting of program synthesis trajectories using GLM-5.2 as the teacher agent. Fine-tuning Qwen3.6-27B on these trajectories increases its ProgramBench average test pass rate from 37.98% to 49.51%, achieving performance comparable to substantially larger frontier models. Moreover, the fine-tuned model consistently improves over the base model across all seven unseen software engineering benchmarks, spanning long-horizon repository generation and translation, bug fixing, feature implementation, and cross-language issue resolution, with absolute gains of 31.00 points on RepoZero-C2Rust, 14.16 on DeepSWE, 10.70/4.56 on NL2Repo-Bench (with/without tests), 5.04 on SWE-bench Verified, 5.93 on SWE-bench Pro, 5.22 on SWE-bench Multilingual, and 4.94 on FeatBench.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!