2607.02344v1 Jul 02, 2026 cs.LG

효율적인 시계열 예측을 위한 자체 게이팅 어텐션

Self-Gating Attention for Efficient Time Series Forecasting

Dezheng Wang
Dezheng Wang
Citations: 86
h-index: 5
Tong Chen
Tong Chen
Citations: 260
h-index: 10
Congyan Chen
Congyan Chen
Citations: 43
h-index: 5
Wei Yuan
Wei Yuan
Citations: 965
h-index: 15
Hongzhi Yin
Hongzhi Yin
Citations: 86
h-index: 5
Shihua Li
Shihua Li
Citations: 316
h-index: 10

트랜스포머 아키텍처는 시계열 예측 분야에서 강력한 잠재력을 보여주며, 멀티 헤드 셀프 어텐션은 과거 시간 간격 간의 시간적 의존성을 파악하는 데 널리 사용됩니다. 그러나 표준 셀프 어텐션은 조회 길이(look-back length)에 대해 이차적인 시간 및 메모리 복잡도를 갖습니다. 이러한 비용은 빠른 추론과 메모리 효율성이 중요한, 자원 제약적이거나 고 처리량의 예측 시스템에서의 활용을 제한할 수 있습니다. 질적 및 양적 분석을 통해 시계열 예측에서 셀프 어텐션 맵이 다양한 시간 간격에 걸쳐 종종 중복 패턴을 포함하고 있음을 관찰했습니다. 이러한 현상은 많은 실제 시계열 데이터에서 반복되는 시간 패턴과 비교적 안정적인 시간 상관관계와 관련될 수 있습니다. 이러한 관찰에 따라, 우리는 입력 의존적인 잔차 구성 요소와 공유 가능한 학습 가능한 행렬로 어텐션 점수를 표현하는 플러그 앤 플레이 어텐션 메커니즘인 Self-Gating Attention (SGA)을 제안합니다. 공유 행렬은 일반적인 어텐션 패턴을 포착하고, 잔차 구성 요소는 입력 의존적인 변형을 포착합니다. 이러한 방식으로 SGA는 표준 어텐션 점수 계산에 사용되는 쿼리 및 키 투영을 피하여 조회 길이에 대한 선형 시간 복잡도와 점수 행렬 메모리 복잡도를 달성합니다. 우리는 SGA를 여러 예측 모델의 백본에 통합하고, 아홉 개의 공개된 실제 데이터 세트(전기, 금융, 기상, 의료 모니터링, 인간 활동 및 기후 기록)에서 표준 셀프 어텐션 및 경량화된 어텐션 변형과 비교했습니다. 결과는 SGA가 공개 벤치마크에서 추론 효율성을 향상시키는 동시에 최첨단 어텐션 메커니즘에 비해 경쟁력 있는 예측 성능을 유지한다는 것을 보여줍니다. 이러한 벤치마크 결과는 실제 배포 환경에서의 적용 가능성을 입증합니다.

Original Abstract

Transformer architectures have shown strong potential in time series forecasting, where multi-head self-attention is widely used to capture temporal dependencies across historical timestamps. However, standard self-attention has quadratic time and memory complexity with respect to the look-back length. This cost may limit its use in resource-constrained or high-throughput forecasting systems, where fast and memory-efficient inference is important. Through qualitative and quantitative analyses, we observe that self-attention maps in time series forecasting often contain redundant patterns across different timestamps. This phenomenon can be related to the repeated temporal patterns and relatively stable temporal correlations in many real-world time series. Motivated by this observation, we propose Self-Gating Attention (SGA), a plug-and-play attention mechanism that represents the attention score with a shared learnable matrix and an input-dependent residual component. The shared matrix captures common attention patterns, while the residual component captures input-dependent variations. In this way, SGA avoids the query and key projections used in standard attention score computation, leading to linear time and score-matrix memory complexity with respect to the look-back length. We integrate SGA into several forecasting backbones and compare it with standard self-attention and lightweight attention variants on nine publicly available real-world datasets covering electricity, finance, weather, medical monitoring, human activity, and climate records. The results show that SGA improves inference efficiency on public benchmarks while maintaining competitive forecasting performance against state-of-the-art attention mechanisms. These benchmark results provide deployment-oriented evidence.

0 Citations
0 Influential
7.5 Altmetric
37.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!