2606.24414v1 Jun 23, 2026 cs.AI

순환 일관성을 갖는 신경망 기반 형식 검증 인증서 설명

Cycle-Consistent Neural Explanation of Formal Verification Certificates

Alberto Pozanco
Alberto Pozanco
Citations: 190
h-index: 8
Daniel Borrajo
Daniel Borrajo
Citations: 48
h-index: 3
A. Rodríguez
A. Rodríguez
Citations: 89
h-index: 6

형식 검증은 시간적 속성의 만족 또는 위반을 증명하는 기계 판독 가능한 인증서를 생성하지만, 이러한 인증서는 비전문가에게는 여전히 불투명합니다. 본 연구에서는 형식이 정확한 자연어 설명을 생성하는 순환 일관성을 갖는 신경망 아키텍처를 제안합니다. 전방 네트워크(NN1)는 인증서를 설명으로 매핑하고, 역방향 네트워크(NN2)는 설명을 통해 인증서를 재구성합니다. 심볼릭 검증기는 루프를 닫고, 차등 가능한 충실도 프록시를 제공합니다. 포인터-제너레이터 메커니즘은 상태 이름을 직접 인증서에서 복사하여 어휘적 정확성을 보장합니다. 우리는 여섯 가지 검증 방법(경계 증명, k-유도, 귀납적 불변량, 라쏘, 도달 가능성, 증거 쌍)을 사용하고, 긍정 및 부정 판별 결과를 모두 포함하는 420개의 테스트 인증서를 대상으로 실험했습니다. 이 데이터는 207개의 명명된 상태를 가진 금융 규정 준수 분야에서 수집되었습니다. 학습된 아키텍처와 하이브리드 추론 시간 라우팅 전략을 결합하면 90.0%의 순환 검증된 정확도를 달성했으며, 이는 네 가지 최첨단 모델 조합 중 가장 좋은 결과를 보인 다중 LLM 기반 방법(76.1%)보다 13.9%p 더 높은 수치입니다. 신경망 모델은 12개의 판별/유형 범주 중 10개에서 우수한 성능을 보였으며, 세 가지 유형에서는 100%의 정확도를 달성했습니다. 제안된 아키텍처는 기존 다중 LLM 기반 방법보다 860배 빠른 추론 속도(인증서당 185ms vs. 160초), 오프라인 작동, 결정론적 출력 및 추론당 비용이 0원이라는 장점을 제공합니다. 이러한 결과는 전문화된 학습 모델이 구조화된 인증서 설명에 대한 일반적인 LLM 프롬프트보다 우수한 성능을 발휘하며, 클라우드 기반 추론의 배포 제약을 해결할 수 있음을 보여줍니다.

Original Abstract

Formal verification produces machine-checkable certificates that attest to the satisfaction or violation of temporal properties, yet these certificates remain opaque to non-specialist stakeholders. We propose a cycle-consistent neural architecture that generates faithful natural language explanations of verification certificates. A forward network NN1 maps certificates to explanations, and an inverse network NN2 reconstructs certificates from explanations; a symbolic verifier closes the loop, providing a differentiable faithfulness proxy. A pointer-generator mechanism ensures lexical grounding by copying state names directly from the certificate. We evaluate on 420 test certificates spanning six verification methods (bounded proof, k-induction, inductive invariant, lasso, reachability, witness pair) in both YES and NO verdict variants, drawn from a financial compliance domain with 207 named states. Our trained architecture, combined with a hybrid inference-time routing strategy, achieves 90.0% cycle-verified soundness, surpassing a multi- LLM few-shot baseline (76.1% for the best of 16 LLM combinations across four frontier models) by 13.9 percentage points. The neural model wins on 10 of 12 verdict/kind categories, with three categories reaching 100% soundness. The architecture offers 860x faster inference (185 ms vs. 160 s per certificate for the full multi-LLM baseline), offline operation, deterministic outputs, and zero per-inference cost. These results demonstrate that trained specialization outperforms general-purpose LLM prompting for structured certificate explanation, while eliminating the deployment constraints of cloud-based inference.

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!