2607.19771v1 Jul 22, 2026 cs.LG

뮤온 기반 동방성 보존 스펙트럴 캡: 이론 및 세 가지 사례 연구

An Isotropy-Preserving Spectral Cap for Muon: Theory and Three Case Studies

Jiachun Li
Jiachun Li
Citations: 12
h-index: 2

뮤온(Muon)과 관련된 행렬 부호 최적화 기법은 대규모 언어 모델의 사전 학습에 점점 더 많이 사용되고 있지만, 이러한 기법이 개별 가중치 행렬의 내부 구조에 미치는 영향은 잘 알려져 있지 않습니다. 본 연구는 단일 이상 가정, 즉 가중치 재조정에 따른 손실 함수의 정확한 스케일 불변성을 기반으로 하는 통합 프레임워크를 제안합니다. 이 가정 하에서, 일반적인 SGD(Stochastic Gradient Descent)는 업데이트 크기에 내재된 1/||W||의 제동 역할을 수행하는 반면, 뮤온의 행렬 부호 단계는 이러한 제동을 제거합니다. 그 결과, Frobenius norm과 spectral norm 모두 더 빠르게 증가하며 (t^{1/2} 대 t^{1/4}), 이는 불안정성을 야기할 수 있습니다. 또한, 스펙트럴 노름의 변화에는 0이 아닌 이차항이 존재한다는 것을 관찰했습니다. 이러한 점을 감안하여, 각 업데이트에서 가장 큰 고유 벡터 방향의 일차 성장 부분만을 제거하는 경량화된 "스펙트럴 캡(spectral cap)"을 사용하여 학습을 중단하지 않고 출력 공분산 W K_X W^T를 제어할 수 있습니다. 가중치는 여전히 다른 방향, 최상위 방향 회전 및 최상위 방향 전환을 통해 학습됩니다. 본 연구에서는 이러한 스펙트럴 캡을 고유값 스펙트럼의 최소 엔트로피(H-infinity)와 관련지었습니다. 그런 다음, 뮤온으로 학습된 세 가지 시스템(nanoGPT 피드포워드 투영, 64 전문가 혼합 전문가 라우터 및 bf16 FlashAttention 블록의 쿼리/키 투영)에 대한 연구를 진행했습니다. 각 경우에서 스펙트럴 캡은 동방성을 증가시키고, 경계 조건 하에서는 특정 시스템(예: 단일 전문가로 붕괴되는 라우터 또는 특정 어텐션 헤드의 거의 발산)의 실패를 방지하는 동시에 검증 손실에는 실질적인 영향을 미치지 않습니다. 본 연구에서 제시된 스케일 불변성 가정은 강력하며, 이러한 소규모 결과는 예비적이라는 점을 강조합니다. 의견 및 제안은 언제든지 환영합니다.

Original Abstract

Muon and related matrix-sign optimizers are increasingly used to pre-train large language models, but their effect on the internal geometry of individual weight matrices is not well understood. This preliminary report proposes a unified framework built on a single idealizing assumption -- exact scale invariance of the loss under weight rescaling, which holds approximately in normalization-heavy networks. Under this assumption, plain SGD carries a built-in 1/||W|| brake on its update size, whereas Muon's matrix-sign step removes that brake, so both the Frobenius and spectral norms drift outward faster (t^{1/2} versus t^{1/4}). We further observe that the spectral-norm perturbation has a non-negative second-order term. This implies that a lightweight "spectral cap" -- which projects out only the first-order growth of the single top singular direction from each update -- can control the output covariance W K_X W^T without freezing training: the weight keeps learning through non-top directions, top-direction rotation, and top switching. We relate this cap to the min-entropy (H-infinity) of the singular-value spectrum. We then study three systems trained with Muon: a nanoGPT feed-forward projection, a 64-expert mixture-of-experts router, and the query/key projections of a bf16 FlashAttention block. In each case the cap increases isotropy and, at the margins -- a router collapsing to a single expert, and the near-divergence of one attention head -- prevents a concrete failure, while leaving validation loss essentially unchanged. We emphasize that the scale-invariance assumption is strong and that these small-scale results are preliminary; comments are welcome.

0 Citations
0 Influential
1 Altmetric
5.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!