2603.09161v1 Mar 10, 2026 cs.LG

잘못된 코드, 올바른 구조: 불완전한 LLM-생성 RTL로부터 학습하는 네틀리스트 표현

Wrong Code, Right Structure: Learning Netlist Representations from Imperfect LLM-Generated RTL

Yi Han
Yi Han
Citations: 104
h-index: 4
Cangyuan Li
Cangyuan Li
Citations: 126
h-index: 5
Siyang Cai
Siyang Cai
Citations: 79
h-index: 5
Ying Wang
Ying Wang
Citations: 83
h-index: 5

효과적인 네틀리스트 표현 학습은 레이블이 지정된 데이터의 부족으로 인해 근본적으로 제약됩니다. 실제 설계는 지적 재산(IP)으로 보호되며, 어노테이션에는 상당한 비용이 소요되기 때문입니다. 기존 연구는 따라서 깨끗한 레이블을 가진 소규모 회로에 집중하여, 현실적인 설계로의 확장성을 제한합니다. 반면, 대규모 언어 모델(LLM)은 대규모 레지스터-전송 레벨(RTL) 코드를 생성할 수 있지만, 기능적 오류로 인해 회로 분석에 사용하기에는 어려움이 있었습니다. 본 연구에서는 중요한 관찰을 합니다. LLM-생성 RTL이 기능적으로 불완전하더라도, 합성된 네틀리스트는 의도된 기능성을 강력하게 나타내는 구조적 패턴을 유지한다는 것입니다. 이러한 통찰력을 바탕으로, 우리는 비용 효율적인 데이터 증강 및 학습 프레임워크를 제안합니다. 이 프레임워크는 LLM-생성 RTL의 불완전성을 체계적으로 활용하여 네틀리스트 표현 학습을 위한 훈련 데이터로 사용하며, 자동 코드 생성부터 다운스트림 작업까지의 엔드-투-엔드 파이프라인을 구축합니다. 서브-회로 경계 식별 및 구성 요소 분류와 같은 회로 기능 이해 작업을 포함한 다양한 규모의 벤치마크를 사용하여 평가를 수행했습니다. 평가 결과, 우리의 노이즈가 있는 합성 데이터셋으로 훈련된 모델은 실제 네틀리스트에 대해 잘 일반화되며, 희소한 고품질 데이터로 훈련된 기존 방법과 동등하거나 더 나은 성능을 보이며, 회로 표현 학습에서의 데이터 병목 현상을 효과적으로 해결합니다.

Original Abstract

Learning effective netlist representations is fundamentally constrained by the scarcity of labeled datasets, as real designs are protected by Intellectual Property (IP) and costly to annotate. Existing work therefore focuses on small-scale circuits with clean labels, limiting scalability to realistic designs. Meanwhile, Large Language Models (LLMs) can generate Register-Transfer-Level (RTL) at scale, but their functional incorrectness has hindered their use in circuit analysis. In this work, we make a key observation: even when LLM-Generated RTL is functionally imperfect, the synthesized netlists still preserve structural patterns that are strongly indicative of the intended functionality. Building on this insight, we propose a cost-effective data augmentation and training framework that systematically exploits imperfect LLM-Generated RTL as training data for netlist representation learning, forming an end-to-end pipeline from automated code generation to downstream tasks. We conduct evaluations on circuit functional understanding tasks, including sub-circuit boundary identification and component classification, across benchmarks of increasing scales, extending the task scope from operator-level to IP-level. The evaluations demonstrate that models trained on our noisy synthetic corpus generalize well to real-world netlists, matching or even surpassing methods trained on scarce high-quality data and effectively breaking the data bottleneck in circuit representation learning.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!