StreamKL: 어텐션 증류를 위한 빠르고 메모리 효율적인 KL 발산 기법
StreamKL: Fast and Memory-Efficient KL Divergence for Boosting Attention Distillation
어텐션 증류는 하나의 어텐션 분포가 다른 분포와 일치하도록 쿨백-라이블러(KL) 발산을 최소화하는 방식으로, 지식 증류, 모델 압축, 지속적 학습 및 희소 어텐션 LLM 학습 등 다양한 분야에서 널리 사용됩니다. 그러나 기존 방법들은 KL 감소를 계산하기 전에 두 가지 어텐션 분포를 모두 메모리에 저장해야 하며, 이로 인해 발생하는 $O(N_QN_K)$의 메모리와 입출력 비용은 긴 컨텍스트 길이에 대해서는 매우 큰 문제가 됩니다. 본 논문에서는 이러한 문제를 해결하기 위해, 어텐션 KL 발산에 대한 최초의 통합 GPU 프라임티브인 StreamKL을 제안합니다. StreamKL은 두 분포 간의 연관된 KL 감소를 위한 새로운 온라인 수식을 도출하여, 쿼리-키 타일을 온칩 SRAM으로 스트리밍하는 단일 패스 포워드 커널을 가능하게 합니다. 역전파 과정에서는 StreamKL이 어텐션 확률을 타일 단위로 재계산하여 이차적인 중간 값을 저장하지 않습니다. 또한 효율적인 GPU 커널을 설계하고 구현했습니다. 실험 결과, StreamKL은 포워드 및 백워드 패스에서 각각 최대 43배와 14배의 속도 향상을 보여주었습니다. 더욱 중요하게는, StreamKL은 어텐션 증류에 필요한 추가 HBM 메모리 공간을 $O(N_QN_K)$에서 $O(1)$로 줄여 단일 GPU에서도 긴 컨텍스트의 증류를 가능하게 합니다.
Attention distillation, which trains one attention distribution to match another by minimizing their Kullback-Leibler (KL) divergence, is widely used in knowledge distillation, model compression, continual learning, and sparse-attention LLM training. However, existing approaches materialize both attention distributions before computing the KL reduction, incurring $O(N_QN_K)$ memory and IO costs that become prohibitive at long context lengths. We present StreamKL, the first fused GPU primitive for attention KL divergence that eliminates this quadratic materialization. StreamKL derives a novel online formulation for the coupled two-distribution KL reduction, enabling a single one-pass forward kernel that streams query-key tiles through on-chip SRAM. For the backward pass, StreamKL recomputes attention probabilities tile-by-tile, avoiding storage of quadratic intermediates. We further design and implement efficient GPU kernels with dedicated optimizations. Experiments show StreamKL delivers up to $43\times$ and $14\times$ speedups over baseline methods in the forward and backward passes, respectively. Most importantly, StreamKL reduces the extra HBM footprint of attention distillation from $O(N_QN_K)$ to $O(1)$, enabling long-context distillation on a single GPU.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.