2604.09967v1 Apr 11, 2026 cs.LG

Muon²: 적응형 2차 모멘트 사전 처리를 통한 뮤온 최적화 성능 향상

Muon$^2$: Boosting Muon via Adaptive Second-Moment Preconditioning

Ruijie Zhang
Ruijie Zhang
Citations: 51
h-index: 4
Yequan Zhao
Yequan Zhao
Citations: 70
h-index: 5
Zhengyang Wang
Zhengyang Wang
Citations: 45
h-index: 3
Yupeng Su
Yupeng Su
Citations: 149
h-index: 4
Zheng Zhang
Zheng Zhang
Citations: 59
h-index: 4
Ziyu Liu
Ziyu Liu
Citations: 30
h-index: 3
Zi Yang
Zi Yang
Citations: 198
h-index: 8

Muon은 신경망 업데이트의 행렬 구조를 활용하여 반복적인 직교화 과정을 통해 대규모 기초 모델 사전 훈련에 유망한 최적화 알고리즘으로 등장했습니다. 그러나 실제 효율성은 최적화 단계마다 여러 번의 뉴턴-슐츠(NS) 반복이 필요하기 때문에 상당한 계산 및 통신 오버헤드가 발생한다는 한계가 있습니다. 본 연구에서는 Muon의 확장 버전인 Muon²를 제안하며, 이는 직교화 전에 Adam 스타일의 적응형 2차 모멘트 사전 처리를 적용합니다. 핵심적인 통찰력은 Muon의 극성 근사 과정에서 가장 큰 어려움이 발생하는 모멘텀 행렬의 불량 조건이며, Muon²는 이 행렬의 스펙트럼을 크게 개선하여 실질적으로 충분한 직교화에 더 빠르게 수렴하도록 합니다. 또한, 우리는 방향 정렬을 통해 실제 직교화 품질을 분석하였으며, Muon²는 각 극성 단계에서 Muon보다 현저한 성능 향상을 보였습니다. 600만 개에서 13억 개 파라미터의 GPT 및 LLaMA 사전 훈련 실험에서, Muon²는 Muon 및 최근 Muon 변형 모델보다 일관되게 우수한 성능을 보였으며, 동시에 NS 반복 횟수를 40% 줄였습니다. 또한, Muon²의 대부분의 이점을 유지하면서 미미한 메모리 오버헤드를 갖는 메모리 효율적인 분해 버전인 Muon²-F를 추가적으로 제안합니다.

Original Abstract

Muon has emerged as a promising optimizer for large-scale foundation model pre-training by exploiting the matrix structure of neural network updates through iterative orthogonalization. However, its practical efficiency is limited by the need for multiple Newton--Schulz (NS) iterations per optimization step, which introduces non-trivial computation and communication overhead. We propose Muon$^2$, an extension of Muon that applies Adam-style adaptive second-moment preconditioning before orthogonalization. Our key insight is that the core challenge of polar approximation in Muon lies in the ill-conditioned momentum matrix, of which the spectrum is substantially improved by Muon$^2$, leading to faster convergence toward a practically sufficient orthogonalization. We further characterize the practical orthogonalization quality via directional alignment, under which Muon$^2$ demonstrates dramatic improvement over Muon at each polar step. Across GPT and LLaMA pre-training experiments from 60M to 1.3B parameters, Muon$^2$ consistently outperforms Muon and recent Muon variants while reducing NS iterations by 40\%. We further introduce Muon$^2$-F, a memory-efficient factorized variant that preserves most of the gains of Muon$^2$ with negligible memory overhead.

8 Citations
0 Influential
4 Altmetric
28.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!