2601.08734v1 Jan 13, 2026 cs.SE

TerraFormer: 정책 기반 검증 피드백을 통해 미세 조정된 LLM을 활용한 자동화된 인프라 코드 생성

TerraFormer: Automated Infrastructure-as-Code with LLMs Fine-Tuned via Policy-Guided Verifier Feedback

Prithwish Jana
Prithwish Jana
Georgia Institute of Technology, Atlanta, USA
Citations: 135
h-index: 7
Sam Davidson
Sam Davidson
Citations: 4
h-index: 1
Bhavana Bhasker
Bhavana Bhasker
Citations: 5
h-index: 1
Andrey Kan
Andrey Kan
Citations: 15
h-index: 1
Anoop Deoras
Anoop Deoras
Citations: 2,990
h-index: 21
Laurent Callot
Laurent Callot
Citations: 103
h-index: 3

인프라 코드(IaC) 자동화는 어려운 과제이며, 대규모 언어 모델(LLM)은 종종 자연어(NL) 입력을 기반으로 잘못된 구성을 생성합니다. 본 논문에서는 TerraFormer라는 신경-기호 프레임워크를 제안합니다. 이 프레임워크는 IaC 생성 및 변환을 위해 지도 학습 기반 미세 조정과 검증기 기반 강화 학습을 결합하며, 형식 검증 도구를 사용하여 구문, 배포 가능성 및 정책 준수에 대한 피드백을 제공합니다. 우리는 다단계 검증 및 반복적인 LLM 자체 수정 과정을 통해 TF-Gen (152,000개 샘플) 및 TF-Mutn (52,000개 샘플)이라는 두 개의 대규모 고품질 NL-to-IaC 데이터 세트를 구축했습니다. Sonnet 3.7, DeepSeek-R1, 및 GPT-4.1과 같이 ~50배 더 큰 모델을 포함한 17개의 최첨단 LLM에 대한 평가 결과, TerraFormer는 IaC-Eval에서 15.94%, TF-Gen (Test)에서 11.65%, TF-Mutn (Test)에서 19.60%의 정확도 향상을 보여주었습니다. TerraFormer는 TF-Gen (Test) 및 TF-Mutn (Test) 모두에서 더 큰 모델보다 뛰어난 성능을 보이며, IaC-Eval에서는 세 번째로 높은 순위를 차지하고, 최적의 모범 사례 및 보안 규정 준수를 달성했습니다.

Original Abstract

Automating Infrastructure-as-Code (IaC) is challenging, and large language models (LLMs) often produce incorrect configurations from natural language (NL). We present TerraFormer, a neuro-symbolic framework for IaC generation and mutation that combines supervised fine-tuning with verifier-guided reinforcement learning, using formal verification tools to provide feedback on syntax, deployability, and policy compliance. We curate two large, high-quality NL-to-IaC datasets, TF-Gen (152k instances) and TF-Mutn (52k instances), via multi-stage verification and iterative LLM self-correction. Evaluations against 17 state-of-the-art LLMs, including ~50x larger models like Sonnet 3.7, DeepSeek-R1, and GPT-4.1, show that TerraFormer improves correctness over its base LLM by 15.94% on IaC-Eval, 11.65% on TF-Gen (Test), and 19.60% on TF-Mutn (Test). It outperforms larger models on both TF-Gen (Test) and TF-Mutn (Test), ranks third on IaC-Eval, and achieves top best-practices and security compliance.

3 Citations
0 Influential
10.5 Altmetric
55.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!