TEON: 텐서 기반 정규화 - 레이어별 뮤온 방식을 넘어선 대규모 언어 모델 사전 훈련 기법
TEON: Tensorized Orthonormalization Beyond Layer-Wise Muon for Large Language Model Pre-Training
뮤온 옵티마이저는 각 레이어별로 행렬 수준의 그래디언트(또는 모멘텀) 정규화를 수행하여 대규모 언어 모델의 사전 훈련에서 뛰어난 성능을 보여주었습니다. 본 연구에서는 뮤온의 일반화된 형태로, 신경망의 그래디언트를 구조화된 고차 텐서로 모델링하여 레이어 간의 정규화를 확장하는 TEON을 제안합니다. 우리는 TEON이 레이어별 뮤온보다 향상된 수렴 보장을 제공하며, 이론적 분석을 바탕으로 실제 구현을 개발하고, 그 효과를 분석했습니다. 제안하는 방법은 130M에서 774M 파라미터 범위를 갖는 GPT 스타일 모델과 60M에서 1B 파라미터 범위를 갖는 LLaMA 스타일 모델을 사용하여 평가했습니다. 실험 결과, TEON은 다양한 모델 크기에서 훈련 및 검증 퍼플렉시티를 꾸준히 개선하며, 다양한 근사 SVD 방식에 대해 높은 안정성을 보였습니다.
The Muon optimizer has demonstrated strong empirical performance in pre-training large language models by performing matrix-level gradient (or momentum) orthogonalization in each layer independently. In this work, we propose TEON, a principled generalization of Muon that extends orthogonalization beyond individual layers by modeling the gradients of a neural network as a structured higher-order tensor. We present TEON's improved convergence guarantee over layer-wise Muon, and further develop a practical instantiation of TEON based on the theoretical analysis with corresponding ablation. We evaluate our approach on two widely adopted architectures: GPT-style models, ranging from 130M to 774M parameters, and LLaMA-style models, ranging from 60M to 1B parameters. Experimental results show that TEON consistently improves training and validation perplexity across model scales and exhibits strong robustness under various approximate SVD schemes.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.