2606.08976v1 Jun 08, 2026 cs.AI

RTL-BenchLS: 대규모 언어 모델을 활용한 RTL 추론 및 생성에 대한 종합 벤치마크

RTL-BenchLS: A Large-Scale Benchmark for RTL Reasoning and Generation with Large Language Models

Shang Liu
Shang Liu
Citations: 618
h-index: 9
Wenji Fang
Wenji Fang
Citations: 713
h-index: 11
Zhiyao Xie
Zhiyao Xie
Citations: 698
h-index: 10
Jing Wang
Jing Wang
Citations: 233
h-index: 5
Yuchao Wu
Yuchao Wu
Citations: 42
h-index: 3
Yugao Zhu
Yugao Zhu
Citations: 8
h-index: 1

대규모 언어 모델(LLM) 기반의 RTL 생성 및 추론은 하드웨어 설계 자동화 분야에서 매우 유망한 연구 방향입니다. 고품질 벤치마크는 이 분야의 발전 상황을 파악하는 데 필수적인 인프라 역할을 합니다. 그러나 기존 RTL 벤치마크는 규모와 작업 범위 측면에서 근본적인 한계를 가지고 있습니다. 일반적으로 다루는 설계는 크기가 작고 단순하며, 작업은 대부분 사양-RTL 생성에 집중되어 있습니다. 현재 최첨단 모델의 성능은 기존 벤치마크에서 이미 최고점에 도달했습니다. 이러한 벤치마크를 확장하는 것은 기본적으로 어렵습니다. 왜냐하면 벤치마킹을 위해서는 사양 및 테스트벤치와 같은 정렬된 레이블이 필요하기 때문입니다. 그러나 실제 설계에 대한 이러한 고품질 데이터는 매우 드물게 존재합니다. 본 논문에서는 위에서 언급한 한계를 해결하는 대규모 벤치마크인 RTL-BenchLS를 소개합니다. RTL-BenchLS는 1만 건 이상의 형식적으로 검증된 Verilog 설계를 포함하며, 기존 벤치마크보다 훨씬 크고 복잡한 설계를 다룹니다. 사양-RTL 생성 외에도, 추론과 생성을 동시에 평가하는 세 가지 새로운 작업을 제안합니다: 양방향 추론, 마스킹 콘텐츠 추론 및 저장소 이슈 해결 작업입니다. 이 중 처음 두 작업은 자기 지도 학습 방식으로 수행되어 확장성의 병목 현상을 직접적으로 해결합니다. 모든 작업은 수동 테스트벤치 없이 형식적 동등성 검증을 통해 확인되었습니다. 우리는 RTL-BenchLS를 사용하여 8개의 LLM을 평가했습니다. 가장 성능이 좋은 모델조차도 자연어 기반의 양방향 추론에서 23%, 마스킹 콘텐츠 추론에서 28%, 저장소 이슈 해결 작업에서 12%의 정확도를 보였습니다. RTL-BenchLS는 기존 벤치마크보다 훨씬 더 높은 난이도를 가지고 있으며, 향후 개선을 위한 충분한 여지를 제공하고 하드웨어 설계에 대한 LLM 기반 방법을 개발하는 데 도움이 되는 지침을 제시합니다.

Original Abstract

LLM-based RTL generation and reasoning is a promising direction for hardware design automation. High-quality benchmarks are critical infrastructure for tracking progress in this direction. However, existing RTL benchmarks face inherent limitations in both scale and task scope. The designs they cover are typically small and simple, and the tasks focus almost entirely on specification-to-RTL generation. Frontier models' performance already saturates on the existing benchmarks. Scaling these benchmarks up is fundamentally difficult because aligned labels are required for benchmarking, such as specifications and testbenches. Such aligned high-quality data are rarely available for real-world designs. We introduce RTL-BenchLS, a large-scale benchmark addressing both limitations above. It contains over 10,000 formally verified Verilog designs, covering substantially larger and more complex designs than existing benchmarks. Beyond specification-to-RTL generation, we propose three novel tasks that jointly evaluate reasoning and generation: round-trip reasoning, masked-content reasoning, and repository-issue reasoning. The first two are self-supervised, which directly resolves the scaling bottleneck. All tasks are verified through formal equivalence checking without any manual testbenches. We evaluate eight LLMs on RTL-BenchLS. Even the best model reaches only 23% on natural-language round-trip reasoning, 28% on masked-content reasoning, and 12% on repository-issue fixing. RTL-BenchLS is substantially more challenging than existing benchmarks. It leaves ample room for future improvement and offers guidance for developing LLM-based methods for hardware design.

0 Citations
0 Influential
5.5 Altmetric
27.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!