2604.10392v1 Apr 12, 2026 cs.LG

추적 가능한 개선을 통한 의도 일치 형식 명세 합성

Intent-aligned Formal Specification Synthesis via Traceable Refinement

Soonho Kong
Soonho Kong
Citations: 2,080
h-index: 15
Zhenyu Liao
Zhenyu Liao
Citations: 2
h-index: 1
Samuel Tenka
Samuel Tenka
Citations: 659
h-index: 2
Udaya Ghai
Udaya Ghai
Citations: 65
h-index: 2
Zhe Ye
Zhe Ye
Citations: 72
h-index: 5
Aidan Z. H. Yang
Aidan Z. H. Yang
Citations: 57
h-index: 2
Huangyuan Su
Huangyuan Su
Citations: 31
h-index: 2
Zhizhen Qin
Zhizhen Qin
Citations: 179
h-index: 7
D. Song
D. Song
Citations: 66
h-index: 3

최근 대규모 언어 모델은 자연어에서 코드를 생성하는 데 점점 더 많이 사용되고 있지만, 정확성을 보장하는 것은 여전히 어려운 과제입니다. 형식 검증은 프로그램이 형식 명세를 만족한다는 것을 증명함으로써 이러한 보장을 얻는 체계적인 방법을 제공합니다. 그러나 실제 코드베이스에서 명세가 자주 누락되는 경우가 많으며, 고품질 명세를 작성하는 것은 비용이 많이 들고 전문 지식이 필요합니다. 본 논문에서는 요구사항 수준의 속성 부여 및 지역적인 수정 기능을 통해 의도에 부합하는 명세를 Lean 언어로 합성하는 추적 가능한 개선 프레임워크인 VeriSpecGen을 제시합니다. VeriSpecGen은 자연어를 원자적인 요구사항으로 분해하고, 생성된 명세를 검증하기 위한 요구사항 중심 테스트를 생성하며, 명시적인 추적성 지도를 제공합니다. 검증에 실패할 경우, 추적성 지도는 실패를 특정 요구사항에 연결하여, 대상 수준의 수정이 가능하도록 합니다. VeriSpecGen은 Claude Opus 4.5를 사용하여 VERINA SpecGen 작업에서 86.6%의 성능을 달성했으며, 다양한 모델 계열 및 규모에서 기준 모델보다 최대 31.8%의 성능 향상을 보였습니다. 추론 시간 성능 향상 외에도, VeriSpecGen 개선 경로로부터 343K개의 학습 예제를 생성했으며, 이러한 예제를 사용하여 학습하면 명세 합성 성능이 62-106% 향상되고 일반적인 추론 능력 향상에도 기여하는 것을 확인했습니다.

Original Abstract

Large language models are increasingly used to generate code from natural language, but ensuring correctness remains challenging. Formal verification offers a principled way to obtain such guarantees by proving that a program satisfies a formal specification. However, specifications are frequently missing in real-world codebases, and writing high-quality specifications remains expensive and expertise-intensive. We present VeriSpecGen, a traceable refinement framework that synthesizes intent-aligned specifications in Lean through requirement-level attribution and localized repair. VeriSpecGen decomposes natural language into atomic requirements and generates requirement-targeted tests with explicit traceability maps to validate generated specifications. When validation fails, traceability maps attribute failures to specific requirements, enabling targeted clause-level repairs. VeriSpecGen achieve 86.6% on VERINA SpecGen task using Claude Opus 4.5, improving over baselines by up to 31.8 points across different model families and scales. Beyond inference-time gains, we generate 343K training examples from VeriSpecGen refinement trajectories and demonstrate that training on these trajectories substantially improves specification synthesis by 62-106% relative and transfers gains to general reasoning abilities.

2 Citations
0 Influential
7.5 Altmetric
39.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!