2606.32007v1 Jun 30, 2026 cs.AI

AxDafny: 에이전트 기반의 Dafny를 이용한 검증된 코드 생성

AxDafny: Agentic Verified Code Generation in Dafny

A. Letson
A. Letson
Citations: 256
h-index: 3
Leopoldo Sarra
Leopoldo Sarra
Citations: 21
h-index: 3
Borja Requena Pozo
Borja Requena Pozo
Citations: 3
h-index: 1
Benjamin Breen
Benjamin Breen
Citations: 187
h-index: 7

본 연구에서는 실행 가능한 코드뿐만 아니라 검증을 위한 증명 자료까지 생성해야 하는 에이전트 기반의 코드 생성 방법을 Dafny에서 다룹니다. 우리는 AxDafny라는 검증기(verifier)가 안내하는 수리 프레임워크를 제시하며, 이를 통해 구현체, 불변식, 단언문 및 종료 논리를 반복적으로 생성합니다. 또한, 250개의 경쟁 프로그래밍 문제를 Dafny로 번역하고 형식적인 사양과 검증기 기반의 평가 환경을 제공하는 벤치마크인 LiveCodeBench-Pro-Dafny (LCB-Pro-Dafny)를 소개합니다. LCB-Pro-Dafny에서 AxDafny는 기존 GPT-5.5 성능에 비해 검증 성공률이 크게 향상되었습니다. 또한, DafnyBench에서 AxDafny는 92.7%의 검증 성공률을 달성하여, 지금까지 보고된 가장 강력한 증명 힌트 기반 시스템보다 6.5% 포인트 더 높은 성능을 보였습니다. 마지막으로, 검증 성공률과 런타임 테스트 성능은 생성된 코드의 서로 다른 측면을 측정한다는 것을 보여줍니다.

Original Abstract

We study agentic code generation in Dafny, where a model must generate both executable code and the proof artifacts for verification. We present AxDafny, a verifier-guided repair framework that iteratively generates implementations, invariants, assertions, and termination arguments. We also introduce LiveCodeBench-Pro-Dafny (LCB-Pro-Dafny), a benchmark of 250 competition-style programming problems translated into Dafny with formal specifications and a verifier-based evaluation harness. On LCB-Pro-Dafny, AxDafny substantially improves verification success over baseline GPT-5.5 performance. On DafnyBench, AxDafny achieves 92.7\% verification success, outperforming the strongest previously reported proof-hint baseline by 6.5 percentage points. Lastly, we show that verification success and runtime test performance measure different aspects of generated code.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!