뮤온 기반 동방성 보존 스펙트럴 캡: 이론 및 세 가지 사례 연구
An Isotropy-Preserving Spectral Cap for Muon: Theory and Three Case Studies
뮤온(Muon)과 관련된 행렬 부호 최적화 기법은 대규모 언어 모델의 사전 학습에 점점 더 많이 사용되고 있지만, 이러한 기법이 개별 가중치 행렬의 내부 구조에 미치는 영향은 잘 알려져 있지 않습니다. 본 연구는 단일 이상 가정, 즉 가중치 재조정에 따른 손실 함수의 정확한 스케일 불변성을 기반으로 하는 통합 프레임워크를 제안합니다. 이 가정 하에서, 일반적인 SGD(Stochastic Gradient Descent)는 업데이트 크기에 내재된 1/||W||의 제동 역할을 수행하는 반면, 뮤온의 행렬 부호 단계는 이러한 제동을 제거합니다. 그 결과, Frobenius norm과 spectral norm 모두 더 빠르게 증가하며 (t^{1/2} 대 t^{1/4}), 이는 불안정성을 야기할 수 있습니다. 또한, 스펙트럴 노름의 변화에는 0이 아닌 이차항이 존재한다는 것을 관찰했습니다. 이러한 점을 감안하여, 각 업데이트에서 가장 큰 고유 벡터 방향의 일차 성장 부분만을 제거하는 경량화된 "스펙트럴 캡(spectral cap)"을 사용하여 학습을 중단하지 않고 출력 공분산 W K_X W^T를 제어할 수 있습니다. 가중치는 여전히 다른 방향, 최상위 방향 회전 및 최상위 방향 전환을 통해 학습됩니다. 본 연구에서는 이러한 스펙트럴 캡을 고유값 스펙트럼의 최소 엔트로피(H-infinity)와 관련지었습니다. 그런 다음, 뮤온으로 학습된 세 가지 시스템(nanoGPT 피드포워드 투영, 64 전문가 혼합 전문가 라우터 및 bf16 FlashAttention 블록의 쿼리/키 투영)에 대한 연구를 진행했습니다. 각 경우에서 스펙트럴 캡은 동방성을 증가시키고, 경계 조건 하에서는 특정 시스템(예: 단일 전문가로 붕괴되는 라우터 또는 특정 어텐션 헤드의 거의 발산)의 실패를 방지하는 동시에 검증 손실에는 실질적인 영향을 미치지 않습니다. 본 연구에서 제시된 스케일 불변성 가정은 강력하며, 이러한 소규모 결과는 예비적이라는 점을 강조합니다. 의견 및 제안은 언제든지 환영합니다.
Muon and related matrix-sign optimizers are increasingly used to pre-train large language models, but their effect on the internal geometry of individual weight matrices is not well understood. This preliminary report proposes a unified framework built on a single idealizing assumption -- exact scale invariance of the loss under weight rescaling, which holds approximately in normalization-heavy networks. Under this assumption, plain SGD carries a built-in 1/||W|| brake on its update size, whereas Muon's matrix-sign step removes that brake, so both the Frobenius and spectral norms drift outward faster (t^{1/2} versus t^{1/4}). We further observe that the spectral-norm perturbation has a non-negative second-order term. This implies that a lightweight "spectral cap" -- which projects out only the first-order growth of the single top singular direction from each update -- can control the output covariance W K_X W^T without freezing training: the weight keeps learning through non-top directions, top-direction rotation, and top switching. We relate this cap to the min-entropy (H-infinity) of the singular-value spectrum. We then study three systems trained with Muon: a nanoGPT feed-forward projection, a 64-expert mixture-of-experts router, and the query/key projections of a bf16 FlashAttention block. In each case the cap increases isotropy and, at the margins -- a router collapsing to a single expert, and the near-divergence of one attention head -- prevents a concrete failure, while leaving validation loss essentially unchanged. We emphasize that the scale-invariance assumption is strong and that these small-scale results are preliminary; comments are welcome.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.