2606.09658v1 Jun 08, 2026 cs.LG

Muon은 Adam보다 더 강력하고 일반화 성능이 뛰어난 특징(feature)을 학습한다

Muon Learns More Robust and Transferable Features than Adam

Fengzhuo Zhang
Fengzhuo Zhang
Citations: 233
h-index: 5
Shihua Zhang
Shihua Zhang
Citations: 23
h-index: 2
Tianyu Ruan
Tianyu Ruan
Citations: 10
h-index: 1
Shuche Wang
Shuche Wang
Citations: 262
h-index: 8

최근 Muon은 대규모 언어 모델(LLM) 및 이미지 분류기의 사전학습에 사용되는 최첨단 최적화 알고리즘으로 떠올랐습니다. Adam 및 SGD보다 효율성이 뛰어나지만, Muon의 특징 학습 능력 우위는 명확하지 않았습니다. 본 논문에서는 Muon의 특징 학습 능력을 강건성(robustness)과 일반화 성능(transferability)이라는 관점에서 분석합니다. 먼저, 사전학습된 모델을 손상된 이미지 및 텍스트 데이터로 평가하여 다양한 아키텍처(트랜스포머 및 CNN 포함)에서 Muon이 Adam 및 SGD보다 일관되게 더 강건한 특징을 학습한다는 것을 보여줍니다. 또한, 학습된 레이어별 분석 도구(probe)를 사용하여 이러한 강건성 우위가 레이어 전체에 걸쳐 더 큰 로짓 마진으로 반영된다는 것을 확인합니다. 둘째, 사전학습된 파라미터를 기반으로 선형 분류기를 훈련하거나 완전한 모델을 미세 조정하여 하위 작업에서 Adam 및 SGD보다 Muon이 학습한 특징이 더욱 효과적으로 일반화된다는 것을 입증합니다. 이러한 일반화 성능 우위는 레이어 간의 숨겨진 상태(hidden state)의 다양성, 즉 효과적인 순위(effective rank)로도 뒷받침됩니다. 마지막으로, 다중 구성 요소 특징을 가진 대표적인 분류 문제에서 Muon은 Adam 및 SGD보다 더 큰 마진과 높은 효과적인 순위를 달성하며, 이는 우리의 실험적 결과에 대한 이론적 근거를 제공합니다.

Original Abstract

Muon has recently emerged as a state-of-the-art optimizer for pretraining Large Language Models (LLMs) and vision classifiers. Despite its efficiency advantage over Adam and SGD, the feature-learning advantage of Muon remains unclear. This paper investigates Muon's feature-learning advantage through the lens of robustness and transferability. First, by evaluating pretrained models on corrupted images and texts, we show that features learned by Muon are consistently more robust than those learned by Adam and SGD across different architectures, including transformers and Convolutional Neural Networks (CNNs). Using trained layer-wise probes, we further show that this robustness advantage is reflected in larger logit margins across layers. Second, by training linear classifiers or fine-tuning full models from pretrained parameters on downstream tasks, we demonstrate that Muon-learned features transfer more effectively than those learned by Adam and SGD. This transferability advantage is further supported by the diversity of hidden states across layers, as measured by effective rank. Finally, in a representative classification problem with multi-component features, we prove that Muon attains larger margins and higher effective rank than Adam and SGD, providing theoretical support for our empirical findings.

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!