AI 에이전트가 실제로 RTL에서 GDS까지 작업을 완료할 수 있을까? 도구 연동 EDA 워크플로우 벤치마킹을 통한 교훈
Can AI Agents Really Complete RTL-to-GDS? Lessons from Benchmarking Tool-Interactive EDA Workflows
LLM 기반 에이전트 시스템은 전자 설계 자동화(EDA) 분야에서 유망한 패러다임으로 부상했으며, 복잡한 설계 워크플로우를 자동화하는 데 강력한 잠재력을 보여주고 있습니다. 그러나 기존의 평가에서는 개별 언어 모델을 독립적인 EDA 작업에 대해 분석하는 경우가 많아, 다양한 에이전트 시스템이 전체 EDA 흐름에서 어떻게 수행되는지에 대한 제한적인 통찰력을 제공합니다. 본 연구에서는 FluxBench를 소개하며, 이는 통일된 프롬프트, 도구 환경 및 기술 라이브러리 설정을 사용하여 AI 에이전트를 엔드투엔드 EDA 워크플로우에 대해 체계적으로 평가하는 것입니다. 우리의 평가는 오픈 소스 툴체인을 사용한 RTL 생성과 산업용 애플리케이션을 위한 상업용 EDA 도구를 사용한 RTL-to-GDS 흐름을 포함한 대표적인 시나리오를 다룹니다. 이러한 워크플로우를 통해 에이전트의 RTL 코드 생성, 반복적 수정, 도구 피드백 활용, 논리 합성, 배치 및 라우팅(P&R), 그리고 공정 변경 주문(ECO) 자동화 능력을 평가합니다. 또한 에이전트 시스템의 효율성을 더욱 자세히 분석하기 위해, 토큰 사용량과 실행 시간에 따른 EDA 결과물의 효과적인 개선 정도를 측정하는 비용 효율성 지표인 Token ROI를 도입했습니다. 실험 결과는 동일한 기반 모델을 사용하는 경우에도, 서로 다른 에이전트 시스템 아키텍처가 최대 86.27%의 성능 차이를 보일 수 있음을 보여줍니다. 또한, 유사한 작업 성능을 보이는 시스템 간에도 Token ROI가 최대 $105.92 imes$까지 다를 수 있습니다. PicoRV32를 사례 연구로 사용한 RTL-to-GDS 흐름에서, FluxEDA는 최대 97.94의 엔드투엔드 점수를 달성하여, 도메인별 EDA 기술을 갖춘 Claude Code보다 최대 $8.39 imes$ 더 높은 성능을 보였습니다. 이러한 결과는 도메인별 전문 지식만으로는 대규모 EDA 시나리오에서 에이전트 성능을 향상시키는 데 충분하지 않으며, 대신 에이전트 시스템 설계와 기반 모델의 기능 모두가 효과적인 자동화된 EDA 워크플로우를 가능하게 하는 데 중요한 역할을 한다는 것을 나타냅니다.
LLM-driven agent systems have emerged as a promising paradigm for electronic design automation (EDA), demonstrating strong potential for automating complex design workflows. However, existing evaluations primarily examine individual language models on isolated EDA tasks, providing limited insight into how different agent systems perform across complete EDA flows. In this work, we present FluxBench, a systematic evaluation of AI agents on end-to-end EDA workflows under unified prompts, tool environments, and technology library settings. Our evaluation covers representative scenarios, including RTL generation with open-source toolchains and an RTL-to-GDS flow using closed-source commercial EDA tools for industrial applications. Through these workflows, we assess agents' capabilities in RTL code generation, iterative repair, tool-feedback utilization, logic synthesis, placement and routing (P&R), and Engineering Change Order (ECO) automation. To further characterize the efficiency of agent systems, we introduce Token ROI, a cost-efficiency metric that measures effective improvements in EDA artifacts relative to token usage and runtime cost. Experimental results show that, even when built on the same foundation model, different agent system architectures can exhibit performance gaps of up to 86.27%. Moreover, among systems with comparable task performance, Token ROI can differ by as much as $105.92\times$. In the RTL-to-GDS flow using PicoRV32 as a case study, FluxEDA achieves an end-to-end score of up to 97.94, outperforming Claude Code equipped with domain-specific EDA skills by up to $8.39\times$. These results indicate that domain-specific skills alone are insufficient to improve agent performance in large-scale EDA scenarios. Instead, both agent system design and foundation model capability play critical roles in enabling effective automated EDA workflows.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.