2607.24027v1 Jul 27, 2026 cs.CV

Sol-Attn: 실시간 어텐션 희소화를 통한 비디오 생성 추론 가속화

Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification

Zeke Xie
Zeke Xie
Citations: 371
h-index: 8
Jincheng Yu
Jincheng Yu
Citations: 1,930
h-index: 8
Junsong Chen
Junsong Chen
Citations: 2,797
h-index: 14
Yitong Li
Yitong Li
Citations: 14
h-index: 2
Song Han
Song Han
Citations: 1,714
h-index: 15
Enze Xie
Enze Xie
Citations: 30,781
h-index: 61
Haopeng Li
Haopeng Li
Citations: 71
h-index: 4
Tian Ye
Tian Ye
Citations: 61
h-index: 5
Haozhe Liu
Haozhe Liu
Citations: 105
h-index: 4
Duomin Wang
Duomin Wang
Citations: 96
h-index: 3
Ruihua Zhang
Ruihua Zhang
Citations: 0
h-index: 0

확고한 화질의 비디오 생성을 위해 디퓨전 트랜스포머가 필수적이지만, 긴 토큰 시퀀스로 인해 어텐션 연산이 주요 병목 현상으로 작용합니다. 훈련 없이 동적으로 어텐션을 희소화하는 방법은 선택된 키-값 블록만 계산함으로써 이러한 병목 현상을 완화하지만, 기존 방법들은 효율성과 정확성 모두에서 어려움을 겪습니다. 그 이유는 다음과 같습니다. (1) 경직되고 예측 불가능하며 비용이 많이 드는 라우팅: 프록시 점수를 기준으로 상위 일정 비율의 블록을 선택하는 방식은 고정된 예산을 사용하지만, 목표 누적 프록시 확률 질량을 달성하기 위해 블록을 유지하는 방식은 동적인 예산을 사용하므로 잠재적으로 불균형이 발생할 수 있으며, 두 가지 모두 프록시 점수를 계산하고 저장하는 데 상당한 오버헤드가 발생합니다. (2) 정보 손실을 초래하는 희소화: 선택되지 않은 블록은 완전히 버려지기 때문에 공격적인 희소화를 사용할 경우 정확도가 저하됩니다. 이러한 한계점을 극복하기 위해, 우리는 비용 효율적인 동적 예산 라우팅 방식을 제안하여 정확도 저하를 최소화합니다. 본 논문에서는 훈련 없이 동작하는 Sol-Attn (Sparsifying online attention)이라는 새로운 방법을 소개합니다. Sol-Attn은 동적 라우팅, 희소 계산 및 근사치 보정을 단일 온라인 소프트맥스 패스 내에서 통합하여 희소 어텐션에서 더 나은 정확도와 효율성의 균형을 제공합니다. Sol-Attn의 핵심은 온라인 소프트맥스 과정에서 프록시 점수를 재사용하여 블록 임계값을 실시간으로 조정하는 것입니다. 이러한 설계는 프록시 맵을 저장하지 않고도 동적이지만 제어 가능한 블록 예산을 가능하게 하며, 동시에 선택되지 않은 블록의 프록시 점수를 직접 재사용하여 그 기여도를 근사합니다. 이미지 및 비디오 생성 작업에 대한 실험 결과, Sol-Attn은 훈련 없는 희소 어텐션의 품질과 효율성 모두를 향상시켜 비디오 생성 및 편집에서 각각 2.1배와 2.3배의 속도 향상을 제공하면서도 시각적 품질을 유지합니다.

Original Abstract

Diffusion transformers are essential for high-fidelity video generation, but long token sequences make attention a dominant inference bottleneck. Training-free dynamic sparse attention alleviates this bottleneck by computing only selected key-value blocks, yet existing methods struggle to sparsify attention both efficiently and accurately for two reasons: (1) Rigid, unpredictable, and costly routing: selecting a fixed fraction of top-ranked blocks by proxy score imposes fixed budgets, whereas retaining blocks to reach a target cumulative proxy probability mass yields dynamic but potentially imbalanced budgets; both incur non-negligible overhead from computing and materializing proxy scores. (2) Lossy keep-or-drop sparsification: unselected blocks are discarded entirely, degrading accuracy under aggressive sparsity. These limitations motivate cheaper dynamic-budget routing while limiting accuracy degradation. In this paper, we introduce training-free Sol-Attn (Sparsifying online attention), which unifies dynamic routing, sparse computation, and approximation correction in a single online-softmax pass, achieving a better accuracy-efficiency trade-off in sparse attention. The core of Sol-Attn is on-the-fly block thresholding with proxy-score reuse, which selects critical blocks by comparing block proxy scores against a threshold during online softmax. This design enables dynamic yet controllable block budgets without materializing the proxy map, while directly reusing the proxy scores of unselected blocks to approximate their contribution. Experiments across image and video generation tasks show that Sol-Attn advances the quality-efficiency frontier of training-free sparse attention, delivering 2.1 times and 2.3 times end-to-end speedups for video generation and editing, respectively, while preserving visual quality.

0 Citations
0 Influential
30 Altmetric
150.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!