Muon²: 적응형 2차 모멘트 사전 처리를 통한 뮤온 최적화 성능 향상
Muon$^2$: Boosting Muon via Adaptive Second-Moment Preconditioning
Muon은 신경망 업데이트의 행렬 구조를 활용하여 반복적인 직교화 과정을 통해 대규모 기초 모델 사전 훈련에 유망한 최적화 알고리즘으로 등장했습니다. 그러나 실제 효율성은 최적화 단계마다 여러 번의 뉴턴-슐츠(NS) 반복이 필요하기 때문에 상당한 계산 및 통신 오버헤드가 발생한다는 한계가 있습니다. 본 연구에서는 Muon의 확장 버전인 Muon²를 제안하며, 이는 직교화 전에 Adam 스타일의 적응형 2차 모멘트 사전 처리를 적용합니다. 핵심적인 통찰력은 Muon의 극성 근사 과정에서 가장 큰 어려움이 발생하는 모멘텀 행렬의 불량 조건이며, Muon²는 이 행렬의 스펙트럼을 크게 개선하여 실질적으로 충분한 직교화에 더 빠르게 수렴하도록 합니다. 또한, 우리는 방향 정렬을 통해 실제 직교화 품질을 분석하였으며, Muon²는 각 극성 단계에서 Muon보다 현저한 성능 향상을 보였습니다. 600만 개에서 13억 개 파라미터의 GPT 및 LLaMA 사전 훈련 실험에서, Muon²는 Muon 및 최근 Muon 변형 모델보다 일관되게 우수한 성능을 보였으며, 동시에 NS 반복 횟수를 40% 줄였습니다. 또한, Muon²의 대부분의 이점을 유지하면서 미미한 메모리 오버헤드를 갖는 메모리 효율적인 분해 버전인 Muon²-F를 추가적으로 제안합니다.
Muon has emerged as a promising optimizer for large-scale foundation model pre-training by exploiting the matrix structure of neural network updates through iterative orthogonalization. However, its practical efficiency is limited by the need for multiple Newton--Schulz (NS) iterations per optimization step, which introduces non-trivial computation and communication overhead. We propose Muon$^2$, an extension of Muon that applies Adam-style adaptive second-moment preconditioning before orthogonalization. Our key insight is that the core challenge of polar approximation in Muon lies in the ill-conditioned momentum matrix, of which the spectrum is substantially improved by Muon$^2$, leading to faster convergence toward a practically sufficient orthogonalization. We further characterize the practical orthogonalization quality via directional alignment, under which Muon$^2$ demonstrates dramatic improvement over Muon at each polar step. Across GPT and LLaMA pre-training experiments from 60M to 1.3B parameters, Muon$^2$ consistently outperforms Muon and recent Muon variants while reducing NS iterations by 40\%. We further introduce Muon$^2$-F, a memory-efficient factorized variant that preserves most of the gains of Muon$^2$ with negligible memory overhead.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.