2608.02502v1 Aug 03, 2026 cs.AI

CMuon: 청킹 기반 모멘텀 직교화 기법을 이용한 디퓨전 트랜스포머 학습 속도 향상 및 안정화

CMuon: Accelerating and Stabilizing Diffusion Transformer Training via Chunked Momentum Orthogonalization

Peng Sun
Peng Sun
Citations: 259
h-index: 6
Chuyan Chen
Chuyan Chen
Citations: 69
h-index: 3
Kun Yuan
Kun Yuan
Citations: 64
h-index: 3

디퓨전 트랜스포머(DiT)는 시각적 생성 모델링 분야에서 최첨단 성능을 달성했지만, 여전히 훈련에 막대한 계산량이 소요됩니다. 최근 제안된 모멘텀 직교화(Muon) 옵티마이저는 AdamW의 유망한 대안으로 제시되었지만, DiT에 직접 적용할 경우 후반 단계에서 최적의 수렴 결과를 얻지 못합니다. 본 논문에서는 이러한 병목 현상의 근본 원인을 규명했습니다. 일반적인 DiT 아키텍처는 계산 효율성을 위해 기능적으로 구별되는 가중치(예: AdaLN 및 QKV 레이어 내의 가중치)를 통합된 텐서로 결합합니다. Muon을 이러한 통합된 텐서에 적용하면 의도치 않게 잠재적인 부분 공간 결합이 발생하여 업데이트 방향을 왜곡하고 전역 최적화를 저해합니다. 이를 해결하기 위해, 본 논문에서는 이러한 행렬을 직교화 전에 독립적인 하위 구성 요소로 분할하는 간단하면서도 매우 효과적인 전략인 Chunked Muon (CMuon)을 제안합니다. 광범위한 실험 결과, CMuon으로 훈련된 6억 7천5백만 개의 파라미터를 가진 DiT 모델이 ImageNet 256에서 단 200 에포크만에 FID 점수 1.18을 달성하는 것으로 나타났습니다. 이는 AdamW를 사용하는 경우보다 2배 이상 빠른 학습 속도를 제공하며, 기존 Muon의 후반 단계 수렴 정체 문제를 효과적으로 해결합니다.

Original Abstract

Diffusion Transformers (DiTs) have achieved state-of-the-art (SOTA) performance in visual generative modeling, yet their training remains computationally prohibitive. While the recently proposed Momentum Orthogonalization (Muon) optimizer offers a promising alternative to AdamW, its direct application to DiTs yields suboptimal late-stage convergence. In this paper, we identify the root cause of this bottleneck: standard DiT architectures fuse functionally distinct weights (e.g., within AdaLN and QKV layers) into unified tensors for computational efficiency. Applying Muon to these fused tensors inadvertently induces implicit subspace coupling, which distorts update directions and degrades global optimization. To address this, we introduce Chunked Muon (CMuon), a simple yet highly effective strategy that partitions these matrices into independent sub-components prior to orthogonalization. Extensive experiments demonstrate that a 675M-parameter DiT trained with CMuon achieves a FID of 1.18 on ImageNet 256 in just 200 epochs. This represents more than a 2x training speedup over AdamW, while effectively overcoming the late-stage convergence plateaus of vanilla Muon.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!