2608.06630v1 Aug 06, 2026 cs.LG

희소성 감지기 (The Sparsity Whisperer)

The Sparsity Whisperer

Dan Alistarh
Dan Alistarh
Citations: 14,718
h-index: 43
Linghao Kong
Linghao Kong
Citations: 9
h-index: 1
Inimai Subramanian
Inimai Subramanian
Citations: 2
h-index: 1
Micah Adler
Micah Adler
Citations: 28
h-index: 2
D. Gutfreund
D. Gutfreund
Citations: 3,294
h-index: 24
N. Shavit
N. Shavit
Citations: 15,228
h-index: 58

가지치기는 대규모 언어 모델의 추론 비용을 줄이지만, 기존 방법들은 주로 큰 활성화 값을 보존하거나 레이어 출력을 재구성하는 데 집중합니다. 우리는 이러한 접근 방식이 특히 희소성에 민감한 MLP 업(up) 및 게이트(gate) 투영에서 수행되는 핵심 연산을 간과한다고 주장합니다. 바로 유사한 입력을 서로 다른 출력으로 분리하는 것입니다. 이는 효과적인 가지치기가 활성화 값뿐만 아니라, 보다 광범위하게 출력 값 사이의 차이를 보존해야 함을 시사합니다. 우리는 이러한 원칙에 기반하여 입력 차이에 대한 정보를 활용하는 가지치기 방법들을 제안합니다. 'Wisp'는 입력 차이 정규화를 사용하여 가중치를 평가하는 1차 업데이트가 필요 없는 방법이며, 'Wisp+'는 각 뉴런이 가장 강력하게 분리하는 입력 쌍을 사용하여 뉴런별로 이 점수를 개선합니다. 마지막으로, 'Whisper'는 약하게 규제된 차분 헤세시안을 재구성 목표로 사용하는 2차 방법입니다. 우리는 70억에서 4050억 개의 파라미터를 가진 Llama 2 및 3.1 모델에 대해 제안하는 방법을 적용한 결과, 우리의 2차 방법은 강력한 재구축 기반 기준보다 일관되게 성능이 향상되었으며, 업데이트가 필요 없는 변형은 활성화 인지 기반 기준보다 성능이 향상되었습니다. 특히 제한된 환경에서 더욱 두드러진 효과를 보였습니다. Wanda 및 SparseGPT와 같은 기존 방법에 비해 제안하는 방법은 구조화된 희소성, 후속 평가 및 다른 모델 패밀리에 대한 개선 효과를 제공합니다. RIA 및 ALPS와 같은 강력한 기술에 우리의 입력 차이에 대한 정보를 활용하는 기준을 추가하면 더 큰 성능 향상을 얻을 수 있으며, 전체 정확도-실행 시간 균형 곡선을 약간의 추가 비용으로 더욱 확장할 수 있습니다. 이러한 결과는 출력 값 사이의 차이를 보존하는 것이 사후 학습된 LLM의 희소화를 위한 광범위하게 유용하고 결합 가능한 신호라는 것을 시사합니다.

Original Abstract

Pruning reduces the inference cost of large language models, but existing criteria primarily preserve large activations or reconstruct layer outputs. We argue that this overlooks a key computation performed by particularly sparsity-sensitive neurons in the MLP up and gate projections: separating similar inputs into dissimilar outputs. This suggests that effective pruning should preserve not only activations, but also the differences between outputs more broadly. We introduce a family of difference-informed pruning methods built upon this principle. Wisp is a first-order, update-free method that scores weights using input-difference norms, and Wisp+ refines this score neuronwise using the input pairs each neuron separates most strongly. Finally, Whisper is a second-order method that uses a lightly regularized difference Hessian as its reconstruction objective. Across Llama 2 and 3.1 models from 7B to 405B parameters, our second-order variant consistently improves over strong reconstruction-based baselines, while our update-free variants improve over activation-aware baselines, especially in constrained settings. The improvements over Wanda and SparseGPT extend to structured sparsity, downstream evaluations, and other model families. Augmenting stronger techniques such as RIA and ALPS with our difference-informed criteria yields further improvements, shifting the overall accuracy-runtime frontier outward at negligible additional cost. These results suggest that preserving output differences is a broadly useful and composable signal for post-training LLM sparsification.

0 Citations
0 Influential
29 Altmetric
145.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!