2605.26895v1 May 26, 2026 cs.LG

크기는 미미하지만 효과는 막대함: 대규모 언어 모델에서의 스케일 벡터에 대하여

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models

Mingze Wang
Mingze Wang
Citations: 331
h-index: 11
Yuxin Fang
Yuxin Fang
Citations: 1,585
h-index: 6
Shuchen Zhu
Shuchen Zhu
Citations: 25
h-index: 4
Binghui Li
Binghui Li
Citations: 13
h-index: 2
Kai Shen
Kai Shen
Citations: 17
h-index: 1
Shu Zhong
Shu Zhong
Citations: 1,032
h-index: 4

최신 대규모 언어 모델(LLM)의 정규화 레이어는 결정적인 정규화 연산과 학습 가능한 스케일 벡터로 구성됩니다. 정규화 연산은 광범위하게 연구되었지만, 널리 사용됨에도 불구하고 스케일 벡터에 대한 이해는 부족합니다. 본 논문에서는 LLM 내의 스케일 벡터를 표현력, 최적화 및 아키텍처 구조라는 관점에서 체계적으로 분석합니다. 첫째, 실험을 통해 스케일 벡터가 모델 파라미터의 극히 작은 부분을 차지하지만, 이를 제거하면 LLM의 사전 학습 성능이 크게 저하된다는 것을 보입니다. 또한, 이론적으로 Pre-Norm 아키텍처에서 스케일 벡터는 표현력을 증가시키지 않지만, 후속 선형 매핑에 대한 자체 증폭 프리컨디셔닝 효과를 통해 최적화를 향상시킨다는 것을 밝힙니다. 둘째, 스케일 벡터에 대한 가중치 감소의 역할을 조사합니다. Input-Norm 및 Output-Norm 레이어를 구분하여 이론적으로 분석한 결과, 가중치 감소는 최적화 및 표현력 측면에서 서로 다른 역할 때문에 전자에 대해서는 유익하지만 후자에는 해롭다는 것을 보여줍니다. 셋째, 이러한 이해를 바탕으로 스케일 벡터에 대한 세 가지 경량이며 상호 보완적인 개선 방안을 제안합니다: 브랜치별 이질성, 선형 매핑 주변의 최적화된 배치, 그리고 크기-방향 재파라미터화. 이론 및 실험 결과는 이러한 각 개선 사항이 일관된 성능 향상을 가져온다는 것을 보여줍니다. 마지막으로, 이러한 개선 사항을 통합하여 하나의 스케일 벡터 전략을 구성하고, 0.12B에서 2B 파라미터 범위의 밀집 모델과 Mixture-of-Experts 모델에 대해 다양한 최적화 알고리즘 및 학습률 스케줄을 사용하여 광범위한 LLM 사전 학습 실험을 통해 평가합니다. 이러한 통합된 전략은 잘 조정된 기준 성능보다 일관되게 낮은 최종 손실 값을 달성하며, 더 우수한 확장성을 보입니다. 또한, 파라미터 및 계산 오버헤드는 미미하게 증가합니다.

Original Abstract

Normalization layers in modern large language models (LLMs) consist of a deterministic normalization operation and a learnable scale vector. While the normalization operation has been extensively studied, the scale vector remains poorly understood despite its ubiquitous use. In this work, we present a systematic study of scale vectors in LLMs from the perspectives of expressivity, optimization, and architectural structure. First, we show empirically that although scale vectors constitute only a negligible fraction of model parameters, removing them substantially degrades LLM pre-training. Our theory further shows that, in Pre-Norm architectures, scale vectors do not increase expressivity; instead, they improve optimization through a self-amplifying preconditioning effect on subsequent linear mappings. Second, we investigate the role of weight decay for scale vectors. By distinguishing Input-Norm and Output-Norm layers, we theoretically show that weight decay is beneficial for the former but harmful for the latter, due to their distinct roles in optimization and expressivity. Third, motivated by this understanding, we propose three lightweight and complementary improvements to scale vectors: branch-specific heterogeneity, improved placement around linear mappings, and magnitude-direction reparameterization. Both theory and experiments show that each improvement yields consistent gains. Finally, we combine these improvements into a unified scale-vector strategy and evaluate it through extensive LLM pre-training experiments on dense and mixture-of-experts models ranging from 0.12B to 2B parameters, across multiple optimizers and learning rate schedules, under industrial-scale token budgets. The unified strategy consistently achieves lower terminal loss than well-tuned baselines and exhibits more favorable scaling behavior, while adding negligible parameter and computational overhead.

1 Citations
0 Influential
5.5 Altmetric
28.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!