2605.14513v1 May 14, 2026 cs.CV

HASTE: 헤드 단위 적응형 희소 어텐션을 이용한 학습 불필요 비디오 디퓨전 가속화

HEART: Exploiting Head Heterogeneity in Sparse Attention for Video Diffusion

Yuexiao Ma
Yuexiao Ma
Citations: 223
h-index: 7
Xuzhe Zheng
Xuzhe Zheng
Citations: 405
h-index: 5
Xiawu Zheng
Xiawu Zheng
Citations: 7,054
h-index: 27
Fei Chao
Fei Chao
Citations: 141
h-index: 4
Rongrong Ji
Rongrong Ji
Citations: 317
h-index: 8
Jing Xu
Jing Xu
Citations: 30
h-index: 1

디퓨전 기반 비디오 생성은 시각적 품질과 시간적 일관성 측면에서 크게 발전했지만, 전체 어텐션의 2차 복잡성으로 인해 실제 적용에는 제한이 있습니다. 학습이 필요 없는 희소 어텐션은 사전 학습된 모델을 재학습하지 않고 가속화할 수 있다는 점에서 매력적이지만, 기존의 온라인 top-p 희소 어텐션은 여전히 마스크 예측에 상당한 비용을 사용하며, 헤드 수준의 이질성을 고려하지 않고 공유된 임계값을 적용합니다. 우리는 이러한 간과된 두 가지 요인이 비디오 DiT에서 학습이 필요 없는 희소 어텐션의 실제적인 속도-품질 균형을 제한한다는 것을 보여줍니다. 이를 해결하기 위해, 우리는 두 가지 모듈형 구성 요소인 헤드 단위 적응형 프레임워크를 제안합니다. 첫째, 쿼리-키 드리프트를 기반으로 불필요한 마스크 예측을 건너뛰는 Temporal Mask Reuse입니다. 둘째, 전역 희소성 예산을 설정하면서 모델 출력 오류를 최소화하여 각 헤드에 최적의 top-p 임계값을 할당하는 Error-guided Budgeted Calibration입니다. Wan2.1-1.3B 및 Wan2.1-14B 모델에서, 우리의 방법은 XAttention 및 SVG2를 지속적으로 개선하여 720P 해상도에서 최대 1.93배의 속도 향상을 달성하면서도 경쟁력 있는 비디오 품질 및 유사성 지표를 유지합니다.

Original Abstract

Sparse attention accelerates video diffusion by allowing each attention head to focus on only a small subset of interactions. Existing methods already construct head-specific sparse patterns conditioned on the input. However, we find that these heads also differ in two less obvious but practically important ways. First, for some heads, sparse attention masks remain stable across many denoising steps, whereas others change rapidly. Second, heads differ substantially in their sensitivity to sparsification: applying the same threshold can induce markedly different errors in the final denoising velocity. Ignoring these differences leads to redundant mask prediction and suboptimal threshold calibration across heads under a global sparsity budget. We present HEART, short for Heterogeneity-Exploiting Adaptive Refresh and Thresholding, a training-free framework that exploits both forms of head heterogeneity. First, Temporal Mask Reuse (TMR) uses a lightweight per-head query-key drift signal to determine whether a cached sparse mask remains reliable across denoising steps, refreshing it only when the drift exceeds a prescribed threshold. Second, Error-guided Budgeted Calibration (EBC) evaluates candidate thresholds offline using a frequency-weighted denoising-velocity error, and assigns each head an appropriate threshold under a global sparsity budget. HEART requires no retraining, weight modification, or sparse-kernel changes, and can be integrated directly into existing sparse-attention pipelines. Across Wan2.1-1.3B, Wan2.1-14B, and HunyuanVideo-13B, HEART consistently pushes the quality--efficiency frontier of advanced sparse attention methods such as XAttention and SVG2 outward.

1 Citations
0 Influential
13.5 Altmetric
68.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!