2605.25612v1 May 25, 2026 cs.LG

활성화 희소성과 평평한 최소값 간의 연결에 대한 연구

Towards the Connection between Activation Sparsity and Flat Minima

Ze Peng
Ze Peng
Citations: 12
h-index: 2
Jian Zhang
Jian Zhang
Citations: 587
h-index: 12
Yang Gao
Yang Gao
Citations: 255
h-index: 8
Yinghuan Shi
Yinghuan Shi
Citations: 6,460
h-index: 37
Lei Qi
Lei Qi
Southeast University
Citations: 3,857
h-index: 29

일반적으로 학습된 트랜스포머 모델의 MLP 블록에서 활성화 희소성이 나타나는 현상은 성능 저하 없이 계산 비용을 크게 줄일 수 있는 기회를 제공합니다. 이 현상을 이론적으로 설명하기 위해 기존 연구들은 활성화 희소성이 데이터 자체의 특성이나 데이터 적합에서 비롯되는 것이 아니라, 학습 과정의 내재적인 편향에서 비롯된다는 것을 보여주었습니다. 그러나 이러한 연결성은 강력한 가정 하에 도출되었으며, 이는 많은 단계를 통해 표준적으로 학습된 심층 모델에는 적용하기 어렵습니다. 본 연구에서는 기존 연구들과 달리, 손실 함수의 평탄성이 MLP 활성화 희소성과 밀접하게 관련되어 있으며, 표준적인 심층 신경망에서 더 약하고 자연스럽게 나타나는 가정으로 작용할 수 있다는 것을 발견했습니다. 구체적으로, 1) MLP 활성화 희소성은 "증강된 평탄성" (평탄성 지표의 가중 합)과 입력 정규화 값 및 MLP의 활성화 그래디언트 곱 간의 비율과 같습니다. 경험적으로 확인한 결과, 이 비율은 학습 과정에서 감소하여 활성화가 희소해지는 현상을 유발합니다. 2) 또한, ReLU 함수 하에서는 활성화 희소성과 동일하지만, 역전파 과정에서의 가지치기를 더욱 용이하게 하고 활성화 희소성보다 안정적인 "미분 희소성"이라는 개념을 제안했습니다. 이론적 결과를 바탕으로, 비율의 분자를 감소시키고 분모를 증가시키는 세 가지 방법을 사용하여 활성화 희소성을 더욱 촉진할 수 있습니다. 이러한 간단하게 적용 가능한 수정 방법은 비율을 효과적으로 줄여 더 많은 활성화를 희소하게 만들 수 있습니다. ImageNet-1K 및 C4 데이터셋에서의 실험 결과, 제안된 방식은 기존 트랜스포머 모델에 비해 추론 시 최소 36%, 학습 시 최소 50%의 성능 향상을 보여주었으며, 이는 추론 및 학습 과정 모두에서 추가적인 비용 절감 가능성을 시사합니다.

Original Abstract

The observation that activation sparsity emerges in MLP blocks of standardly trained Transformers offers an opportunity to drastically reduce computation costs without sacrificing performance. To theoretically explain this phenomenon, existing works have shown that activation sparsity does not result from the data properties or data fitting but from the implicit bias of the training process. However, these connections are obtained with strong assumptions, which cannot be applied to deep models standardly trained with a large number of steps. Different from these works, we find that the flatness of loss landscapes is also closely related to the MLP activation sparsity and can serve as a weaker and naturally emerging assumption standard deep networks. Specifically, we find that 1) the MLP activation sparsity equals a ratio between "augmented flatness" (a weighted sum of flatness measures) and the product of the input norm and activation gradient of the MLP. We empirically find that this ratio decreases during training, leading to sparse activations. 2) We also propose the notion of derivative sparsity, which reduces to activation sparsity under ReLU, but further enables pruning in the backward propagation and is more stable than activation sparsity. With the theoretical findings, we can further encourage activation sparsity by decreasing the numerator and increasing the denominator of the ratio using three methods. These plug-and-play modifications can effectively reduce the ratio and produce sparser activations. Experiments on ImageNet-1K and C4 demonstrate relative improvements of at least 36% on inference sparsity and at least 50% on training sparsity over vanilla Transformers, indicating further potential cost reduction in both inference and training

0 Citations
0 Influential
18.5 Altmetric
92.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!