2602.11016v1 Feb 11, 2026 cs.AR

버퍼에서 레지스터로: 하이브리드 본딩 3D NPU 공동 설계 기반의 정밀한 FlashAttention 구현

From Buffers to Registers: Unlocking Fine-Grained FlashAttention with Hybrid-Bonded 3D NPU Co-Design

Jinxin Yu
Jinxin Yu
Citations: 14
h-index: 1
Yudong Pan
Yudong Pan
Citations: 92
h-index: 3
Mengdi Wang
Mengdi Wang
Citations: 224
h-index: 8
Huawei Li
Huawei Li
Citations: 49
h-index: 5
Yi Han
Yi Han
Citations: 104
h-index: 4
Xiaowei Li
Xiaowei Li
Citations: 197
h-index: 8
Ying Wang
Ying Wang
Citations: 172
h-index: 5

Transformer 기반 모델은 현대 AI 워크로드에서 지배적인 위치를 차지하지만, 2차원 주의 메커니즘의 복잡성과 끊임없이 증가하는 모델 크기로 인해 메모리 병목 현상을 심화시킵니다. Groq 및 Cerebras와 같은 기존 가속기는 대용량 온칩 캐시를 사용하여 오프칩 트래픽을 줄이는 반면, FlashAttention과 같은 알고리즘 혁신은 큰 주의 행렬을 구체화하지 않도록 연산자를 결합합니다. 그러나 오프칩 트래픽이 감소함에 따라, 우리의 측정 결과에 따르면 긴 시퀀스 워크로드에서 온칩 SRAM 접근이 전체 에너지 소비의 60% 이상을 차지하며, 이는 캐시 접근을 새로운 병목 지점으로 만듭니다. 우리는 수직으로 분할된 PE 티어 간의 레지스터-레지스터 통신을 가능하게 하는 하이브리드 본딩 3D 스택 공간 가속기인 3D-Flow를 제안합니다. 2D 멀티 어레이 아키텍처는 NoC 기반 라우터-투-라우터 전송에 의해 제한되는 반면, 3D-Flow는 10um 미만의 수직 TSV를 활용하여 최소한의 오버헤드로 사이클 단위의 연산자 파이프라이닝을 유지합니다. 이 아키텍처를 기반으로, 우리는 각 티어 간의 지연 시간을 균형 있게 조정하여 온칩 SRAM 왕복 없이 버블 없는 수직 데이터 흐름을 형성하는 정밀한 스케줄링 방법인 3D-FlashAttention을 설계했습니다. Transformer 워크로드(OPT 및 QWEN 모델)에 대한 평가 결과, 우리의 3D 공간 가속기는 46~93%의 에너지 소비 감소와 최첨단 2D 및 3D 설계에 비해 1.4배에서 7.6배의 속도 향상을 달성했습니다.

Original Abstract

Transformer-based models dominate modern AI workloads but exacerbate memory bottlenecks due to their quadratic attention complexity and ever-growing model sizes. Existing accelerators, such as Groq and Cerebras, mitigate off-chip traffic with large on-chip caches, while algorithmic innovations such as FlashAttention fuse operators to avoid materializing large attention matrices. However, as off-chip traffic decreases, our measurements show that on-chip SRAM accesses account for over 60% of energy in long-sequence workloads, making cache access the new bottleneck. We propose 3D-Flow, a hybrid-bonded, 3D-stacked spatial accelerator that enables register-to-register communication across vertically partitioned PE tiers. Unlike 2D multi-array architectures limited by NoC-based router-to-router transfers, 3D-Flow leverages sub-10 um vertical TSVs to sustain cycle-level operator pipelining with minimal overhead. On top of this architecture, we design 3D-FlashAttention, a fine-grained scheduling method that balances latency across tiers, forming a bubble-free vertical dataflow without on-chip SRAM roundtrips. Evaluations on Transformer workloads (OPT and QWEN models) show that our 3D spatial accelerator reduces 46-93% energy consumption and achieves 1.4x-7.6x speedups compared to state-of-the-art 2D and 3D designs.

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!