AgenticCANN: 지식 기반 에이전트 진화를 통한 자동화된 Ascend C 연산자 생성
AgenticCANN: Automated Ascend C Operator Generation via Knowledge-Augmented Agentic Evolution
Ascend C 연산자 최적화는 NPU(신경 처리 장치)의 추론 성능에 매우 중요하지만, 깊이 있는 하드웨어 전문 지식이 필요합니다. 대규모 언어 모델(LLM)은 자동화된 CUDA 커널 생성에서 가능성을 보여주었지만, Ascend C의 근본적으로 다른 프로그래밍 모델은 아직 탐구되지 않은 고유한 과제를 제시합니다. 본 논문에서는 지식 기반 에이전트 진화 프레임워크인 AgenticCANN을 제안하며, 이는 저자원 NPU 환경에서 자동화된 Ascend C 연산자 합성을 위해 특별히 설계되었습니다. 낯설지 않은 하드웨어에 대한 심각한 플랫폼 지식 부족 문제를 해결하기 위해, AgenticCANN은 구조적이고 다단계의 도메인 인사이트를 개발 라이프사이클 전반에 걸쳐 제공하는 지식 기반 생성 시스템을 통합하여 상위 단계의 실현 가능성 병목 현상을 해소합니다. 이 기반 위에, AgenticCANN은 LLM과의 상호 작용 모드를 특정 생성 및 진화 단계에 맞춰 동적으로 조정하는 스테이지 적응형 에이전트 진화 전략을 특징으로 하며, 높은 탐색적 후보 발견과 높은 수렴 성능 튜닝 사이의 균형을 유지합니다. 화웨이 Ascend 910B에서 수행된 광범위한 실험 결과, 제안하는 방법은 다섯 가지 패턴 범주에 속하는 여섯 개 연산자에 대해 요소별 및 정규화 연산자의 경우 90~100%의 실현 가능성을 달성하고, 융합 연산자의 경우 56%의 실현 가능성을 달성하며, 최대 6.65배의 추론 속도 향상을 보였습니다. 추가 분석 결과, 지식 주입은 요소별 연산자의 경우 57%에서 86%로 실현 가능성을 단조롭게 개선하는 것으로 나타났으며, 이는 특정 연산자가 아닌 일반적인 이점을 제공함을 보여줍니다.
Ascend C operator optimization is critical for NPU (Neural Processing Unit) inference performance but requires deep hardware expertise.While large language models (LLMs) have shown promise in automated CUDA kernel generation, the fundamentally different programming model of Ascend C introduces unique challenges that remain unexplored. In this paper, we propose AgenticCANN, a knowledge-augmented agentic evolution framework specifically tailored for automated Ascend C operator synthesis in low-corpus NPU environments.To overcome the severe platform knowledge deficit on unfamiliar hardware, AgenticCANN incorporates a knowledge-orchestrated generation system that delivers structured, multi-level domain insights across the development lifecycle to resolve the upstream feasibility bottleneck.Building on this foundation, it features a stage-adaptive agentic evolution strategy that dynamically aligns LLM interaction modes with specific generation and evolution phases, balancing high-exploration candidate discovery with high-convergence performance tuning.Extensive experiments on Huawei Ascend 910B across six operators spanning five pattern categories demonstrate that our method achieves 90 to 100 percent feasibility on elementwise and normalization operators, 56% on fusion operators, and up to 6.65$\times$ speedup on 1B Pangu model inference kernels. Further analysis reveals that knowledge injection monotonically improves feasibility from 57% to 86% on elementwise operators, demonstrating its general rather than operator-specific benefit.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.