버퍼에서 레지스터로: 하이브리드 본딩 3D NPU 공동 설계 기반의 정밀한 FlashAttention 구현
From Buffers to Registers: Unlocking Fine-Grained FlashAttention with Hybrid-Bonded 3D NPU Co-Design
Transformer 기반 모델은 현대 AI 워크로드에서 지배적인 위치를 차지하지만, 2차원 주의 메커니즘의 복잡성과 끊임없이 증가하는 모델 크기로 인해 메모리 병목 현상을 심화시킵니다. Groq 및 Cerebras와 같은 기존 가속기는 대용량 온칩 캐시를 사용하여 오프칩 트래픽을 줄이는 반면, FlashAttention과 같은 알고리즘 혁신은 큰 주의 행렬을 구체화하지 않도록 연산자를 결합합니다. 그러나 오프칩 트래픽이 감소함에 따라, 우리의 측정 결과에 따르면 긴 시퀀스 워크로드에서 온칩 SRAM 접근이 전체 에너지 소비의 60% 이상을 차지하며, 이는 캐시 접근을 새로운 병목 지점으로 만듭니다. 우리는 수직으로 분할된 PE 티어 간의 레지스터-레지스터 통신을 가능하게 하는 하이브리드 본딩 3D 스택 공간 가속기인 3D-Flow를 제안합니다. 2D 멀티 어레이 아키텍처는 NoC 기반 라우터-투-라우터 전송에 의해 제한되는 반면, 3D-Flow는 10um 미만의 수직 TSV를 활용하여 최소한의 오버헤드로 사이클 단위의 연산자 파이프라이닝을 유지합니다. 이 아키텍처를 기반으로, 우리는 각 티어 간의 지연 시간을 균형 있게 조정하여 온칩 SRAM 왕복 없이 버블 없는 수직 데이터 흐름을 형성하는 정밀한 스케줄링 방법인 3D-FlashAttention을 설계했습니다. Transformer 워크로드(OPT 및 QWEN 모델)에 대한 평가 결과, 우리의 3D 공간 가속기는 46~93%의 에너지 소비 감소와 최첨단 2D 및 3D 설계에 비해 1.4배에서 7.6배의 속도 향상을 달성했습니다.
Transformer-based models dominate modern AI workloads but exacerbate memory bottlenecks due to their quadratic attention complexity and ever-growing model sizes. Existing accelerators, such as Groq and Cerebras, mitigate off-chip traffic with large on-chip caches, while algorithmic innovations such as FlashAttention fuse operators to avoid materializing large attention matrices. However, as off-chip traffic decreases, our measurements show that on-chip SRAM accesses account for over 60% of energy in long-sequence workloads, making cache access the new bottleneck. We propose 3D-Flow, a hybrid-bonded, 3D-stacked spatial accelerator that enables register-to-register communication across vertically partitioned PE tiers. Unlike 2D multi-array architectures limited by NoC-based router-to-router transfers, 3D-Flow leverages sub-10 um vertical TSVs to sustain cycle-level operator pipelining with minimal overhead. On top of this architecture, we design 3D-FlashAttention, a fine-grained scheduling method that balances latency across tiers, forming a bubble-free vertical dataflow without on-chip SRAM roundtrips. Evaluations on Transformer workloads (OPT and QWEN models) show that our 3D spatial accelerator reduces 46-93% energy consumption and achieves 1.4x-7.6x speedups compared to state-of-the-art 2D and 3D designs.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.