2603.04127v1 Mar 04, 2026 cs.LG

데이터 기반 랜덤 특징 커널을 활용한 트랜스포머

Data-Aware Random Feature Kernel for Transformers

Amirhossein Farzam
Amirhossein Farzam
Citations: 14
h-index: 2
Hossein Mobahi
Hossein Mobahi
Citations: 122
h-index: 4
N. Miller
N. Miller
Citations: 172
h-index: 6
Luke Sernau
Luke Sernau
Citations: 21
h-index: 2

트랜스포머는 다양한 분야에서 뛰어난 성능을 보이지만, 2차 복잡도를 가지는 어텐션 메커니즘은 확장성을 제한하는 요인입니다. Performers에서 사용된 랜덤 특징 어텐션은 양의 랜덤 특징을 이용하여 소프트맥스 커널을 근사함으로써 어텐션 복잡도를 시퀀스 길이에 선형적으로 줄일 수 있습니다. 그러나 사전 학습된 모델에서 쿼리와 키는 일반적으로 등방성이 아닙니다. 이는 등방성 샘플링 방식에서 높은 몬테 카를로 분산을 유발하며, 모델을 재학습하거나 큰 특징 버짓을 사용해야 합니다. 중요 샘플링은 입력 데이터의 기하학적 특성에 맞춰 샘플링 분포를 조정하여 이 문제를 해결할 수 있지만, 복잡한 데이터 의존적인 제안 분포는 종종 계산하기 어렵습니다. 본 연구에서는 소프트맥스 커널을 데이터에 맞게 정렬함으로써, 중요 샘플링을 위한 계산 가능한 최소 분산 제안 분포를 제공하고, 더 나은 학습 안정성을 보이는 어텐션 메커니즘을 얻을 수 있음을 보여줍니다. 이러한 발견에 기반하여, 데이터 정렬된 커널 기하학 구조를 특징으로 하는 DARKFormer라는 데이터 기반 랜덤 특징 커널 트랜스포머를 제안합니다. DARKFormer는 랜덤 투영의 공분산을 학습하여, 데이터 정렬된 커널에 대한 중요 샘플링된 양의 랜덤 특징 추정기를 효율적으로 구현합니다. 실험 결과, DARKFormer는 정확한 소프트맥스 어텐션과의 성능 격차를 줄이며, 특히 사전 학습된 표현이 등방성이 아닌 파인튜닝 환경에서 더욱 효과적입니다. DARKFormer는 랜덤 특징의 효율성과 데이터 기반 커널을 결합하여, 자원 제약적인 환경에서 커널 기반 어텐션을 발전시킵니다.

Original Abstract

Transformers excel across domains, yet their quadratic attention complexity poses a barrier to scaling. Random-feature attention, as in Performers, can reduce this cost to linear in the sequence length by approximating the softmax kernel with positive random features drawn from an isotropic distribution. In pretrained models, however, queries and keys are typically anisotropic. This induces high Monte Carlo variance in isotropic sampling schemes unless one retrains the model or uses a large feature budget. Importance sampling can address this by adapting the sampling distribution to the input geometry, but complex data-dependent proposal distributions are often intractable. We show that by data aligning the softmax kernel, we obtain an attention mechanism which can both admit a tractable minimal-variance proposal distribution for importance sampling, and exhibits better training stability. Motivated by this finding, we introduce DARKFormer, a Data-Aware Random-feature Kernel transformer that features a data-aligned kernel geometry. DARKFormer learns the random-projection covariance, efficiently realizing an importance-sampled positive random-feature estimator for its data-aligned kernel. Empirically, DARKFormer narrows the performance gap with exact softmax attention, particularly in finetuning regimes where pretrained representations are anisotropic. By combining random-feature efficiency with data-aware kernels, DARKFormer advances kernel-based attention in resource-constrained settings.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!