2607.25291v1 Jul 28, 2026 cs.CL

CoSA: 프록시-커널 공동 설계 기반의 희소 어텐션을 통한 장문 컨텍스트 추론 가속화

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention

Yufei Xue
Yufei Xue
Citations: 2,257
h-index: 7
Guanghua Yu
Guanghua Yu
Citations: 50
h-index: 4
Hong Liu
Hong Liu
Citations: 36
h-index: 2
Lin Niu
Lin Niu
Citations: 456
h-index: 4
Jun Zhang
Jun Zhang
Citations: 34
h-index: 4
Hanyong Shao
Hanyong Shao
Citations: 50
h-index: 5
Jianchen Zhu
Jianchen Zhu
Citations: 479
h-index: 7
Siran Liu
Siran Liu
Citations: 46
h-index: 1
Wei Liu
Wei Liu
Citations: 0
h-index: 0

자가 어텐션의 제곱에 비례하는 계산 비용은 장문 컨텍스트 추론을 매우 비효율적으로 만듭니다. 프록시 기반 블록 희소 어텐션은 이러한 문제를 해결하기 위한 실용적인 방법으로 떠올랐습니다. 기존 방식은 일반적으로 프록시를 사용하여 이진 희소 마스크를 예측하고, 커널이 이 마스크를 활용하여 희소 어텐션 계산을 수행합니다. 이러한 접근 방식은 제한된 예산 하에서는 효과적이지만, 예산이 더욱 줄어들면 추정된 프록시는 필연적으로 중요한 블록들을 누락하게 되고, 커널은 희소 마스크를 기계적으로 적용할 수밖에 없어 모델 정확도가 현저히 떨어지는 문제가 발생합니다. 본 논문에서는 커널 인식 프록시(KAP)와 정렬 건너뛰기 커널(OSK)을 결합한 두 단계의 학습이 필요 없는 희소 어텐션, 즉 CoSA를 제안합니다. 첫 번째 단계에서 KAP는 적절한 예산 하에 블록을 선택하고, 커널 내부 루프에서 KV 페이지를 방문하는 순서를 결정하는 정렬된 마스크를 생성합니다. 두 번째 단계에서는 OSK가 이 마스크를 적용하고, 온라인 소프트맥스 통계 정보를 바탕으로 더욱 많은 블록을 건너뛰어 제한적인 예산 하에서도 효율성을 높입니다. 다양한 LLM 기반 모델과 장문 컨텍스트 벤치마크에서 CoSA는 더 낮은 예산으로도 높은 정확도를 달성합니다. 주목할 만한 점은 CoSA가 컨텍스트 길이가 128K인 환경에서 어텐션 속도를 4.93배 향상시키고, 첫 번째 토큰 생성까지의 시간을 2.53배 단축하며, 성능 저하 없이 이러한 효율성을 달성했다는 것입니다.

Original Abstract

The quadratic cost of self-attention makes long-context inference prohibitively expensive, and proxy-based block-sparse attention has become a practical remedy. Existing methods typically rely on a proxy to predict a binary sparse mask and a kernel to consume this mask and perform sparse attention computation. Such an approach is effective under moderate budgets. However, as the budget tightens, the estimated proxy inevitably drops some salient blocks, while the kernel can only apply the sparse mask mechanically, leading to an evident drop in model accuracy. We propose CoSA, a two-stage training-free Sparse Attention under proxy-kernel CO-design, which couples a Kernel-Aware Proxy (KAP) with an Ordered-Skipping Kernel (OSK). In the first stage, the KAP selects blocks under a moderate budget and produces an ordered mask that prescribes the order in which KV pages are visited in the kernel inner loop. In the second stage, the OSK applies this mask and skips more blocks under a tightened budget given online-softmax statistics. Across mainstream LLM backbones and long-context benchmarks, CoSA attains higher accuracy at lower budgets. Impressively, CoSA achieves a 4.93$\times$ attention speedup and reduces end-to-end Time-to-First-Token by 2.53$\times$ under a context length of 128K with negligible performance degradation. Code is available at https://github.com/Tencent/AngelSlim.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!