2607.29674v1 Jul 31, 2026 math.OC

Muon을 위한 부호 압축: SignMuon, MuonSign 및 오류 피드백의 한계

Sign compression for Muon: SignMuon, MuonSign, and the Limits of Error Feedback

Alexey Kravatskiy
Alexey Kravatskiy
Citations: 3
h-index: 1
Maria Smirnova
Maria Smirnova
Citations: 0
h-index: 0

SignMuon은 각 파라미터당 1비트만 사용하여 Muon 업데이트를 압축하며, 이는 행렬 정보를 활용하는 최적화 알고리즘을 극히 낮은 통신 예산 하에서 실행하는 가장 직접적인 방법입니다. SignMuon은 실제 환경에서 SignSGD보다 성능이 우수하지만, 선형 함수에 대해서도 수렴할 수 있습니다. 선형 최소화 오라클(LMO) 이후가 아닌, 그 전에 그래디언트에 부호를 부여하는 것으로는 이러한 문제를 해결할 수 없습니다. 우리는 sign-before (MuonUSign) 및 sign-on-both-sides (MuonSign) 모두 수렴하지 않는 작은 예시를 구성했습니다. 즉, 오라클 주변에 부호를 배치하더라도 일반적으로 수렴이 보장되지 않습니다. 편향된 압축기를 수정하는 일반적인 방법인 오류 피드백은 SignMuon을 구원할 수 없습니다. 오류 피드백을 Muon의 출력에 적용하면 모든 평활도 상수, 스텝 크기 및 모멘텀에서 실패할 수 있습니다. 반면, 그래디언트에 오류 피드백을 적용하면 EF21-MuonUSign과 EF21-MuonSign은 평활한 비볼록 문제에서 표준적인 $\mathcal{O}(T^{-1/2})$ 수렴률을 달성하며, 특히 후자는 각 방향으로 1비트를 사용합니다. 실험 결과, 중앙 집중식 CIFAR-10, 분산형 CIFAR-10 및 nanoGPT 속도 테스트에서 가장 강력한 압축 방법은 LMO 이후에 부호를 부여하는 방식이며, 이는 우리가 수렴하지 않는다고 증명한 방식입니다. 반면, 수렴이 보장되는 변형들은 그 뒤를 잇습니다. 이러한 규모에서는 보장된 성능보다 휴리스틱하게 압축을 수행하는 것이 더 중요합니다.

Original Abstract

SignMuon compresses the Muon update to one bit per parameter by taking its elementwise sign, providing the most direct way to run a matrix-aware optimizer under an extremely low communication budget. It outperforms SignSGD in practice, yet it can ascend even on a linear function. Signing the gradient before the Linear Minimization Oracle (LMO), rather than after, does not repair this: we construct a small explicit instance on which sign-before (MuonUSign) and sign-on-both-sides (MuonSign) ascend as well, so no placement of the sign around the oracle descends in general. Error feedback, the standard remedy for a biased compressor, does not rescue SignMuon: when applied to Muon's output, error feedback can fail for every smoothness constant, step size, and momentum. Applied to the gradient, error feedback does work, and EF21-MuonUSign and EF21-MuonSign attain the standard $\mathcal{O}(T^{-1/2})$ rate for the squared gradient norm on smooth nonconvex problems, the latter at one bit in each direction. Experiments then reverse the ordering: across centralized CIFAR-10, federated CIFAR-10, and the nanoGPT speedrun, the strongest compressed method is consistently sign-after-the-LMO, precisely the placement we prove divergent, with the provably convergent variants trailing it. Compressing after the LMO, a heuristic, matters more at these scales than the guarantee does.

0 Citations
0 Influential
0.5 Altmetric
2.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!