2608.03457v1 Aug 04, 2026 cs.AI

LLaDA MoE v2: Mixture-of-Experts 확산 언어 모델의 확장 연구

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Zhenzhong Lan
Zhenzhong Lan
Citations: 247
h-index: 7
Jianguo Li
Jianguo Li
Citations: 285
h-index: 7
Huabin Liu
Huabin Liu
Citations: 142
h-index: 3
Yipeng Xing
Yipeng Xing
Citations: 170
h-index: 5
Xiaolu Zhang
Xiaolu Zhang
Citations: 1,075
h-index: 5
Jingyang Ou
Jingyang Ou
Citations: 1,065
h-index: 4
Fengqi Zhu
Fengqi Zhu
Citations: 1,468
h-index: 6
Zebin You
Zebin You
Citations: 962
h-index: 6
Wayne Xin Zhao
Wayne Xin Zhao
Citations: 194
h-index: 8
Jirong Wen
Jirong Wen
Citations: 1,187
h-index: 9

확산 언어 모델(dLLM)은 오토리거시브(AR) 언어 모델에 대한 대안을 제공하지만, Mixture-of-Experts (MoE) dLLM의 확장 특성에 대한 이해는 아직 부족합니다. 본 연구에서는 최적화 하이퍼파라미터, 컴퓨팅 할당 및 아키텍처가 MoE dLLM에서 어떻게 확장되는지 체계적으로 분석하고, AR 모델에 대해 이전 보고된 확장 추세와 양적인 차이를 확인했습니다. 특히, 최적화 측면에서 이상적인 배치 크기는 더 빠르게 증가하는 반면, 이상적인 학습률은 컴퓨팅 용량 증가에 따라 더욱 빠르게 감소합니다. 모델-데이터 할당 측면에서는 IsoFLOP 분석을 통해 데이터 측면의 경향이 약간 나타나며, 활성화된 모델 측 컴퓨팅보다 토큰 예산이 더 빠르게 증가합니다. MoE 아키텍처 측면에서, 확장 규모가 커질수록 고정된 활성화 용량에서 더 큰 전문가 풀을 선호하는 경향이 있으며, 적당한 수준의 전문가 세분화는 꾸준히 효과적이며, 공유 전문가에 할당되는 활성화 용량 비율은 확장 규모에 관계없이 안정적으로 유지됩니다. 이러한 연구 결과를 바탕으로, 30B-A3B 크기의 dLLM인 LLaDA MoE v2를 처음부터 235억 개의 토큰으로 학습했습니다. Qwen3의 사전 학습 토큰 수보다 약 65% 적은 데이터로도, LLaDA MoE v2는 여러 지식, 추론 및 코딩 벤치마크에서 Qwen3에 근접하는 성능을 보였습니다. 지도 학습 미세 조정만으로도, SDAR Chat보다 8개의 추론 및 코딩 벤치마크 중 7개에서 더 나은 성능을 보였으며, 여러 작업에서 Qwen3와 비슷한 수준의 결과를 얻었습니다. 이러한 결과는 MoE dLLM에 대한 실질적인 확장 법칙과 설계 원칙을 제시합니다.

Original Abstract

Diffusion language models (dLLMs) offer an alternative to autoregressive (AR) language modeling, yet the scaling behavior of Mixture-of-Experts (MoE) dLLMs remains poorly understood. We systematically characterize how optimization hyperparameters, compute allocation, and architecture scale for MoE dLLMs, identifying quantitative differences from scaling trends previously reported for AR models. Specifically, for optimization, the optimal nominal batch size grows faster, while the optimal learning rate decays more rapidly with compute. For model--data allocation, IsoFLOP analysis reveals a slight data-side tilt: the optimal token budget grows faster than activated model-side computation. For MoE architecture, larger scales increasingly favor larger expert pools at fixed activated capacity, while moderate expert granularity remains consistently effective and the preferred fraction of activated capacity assigned to shared experts remains stable across scales. Guided by these findings, we train LLaDA MoE v2, a 30B-A3B dLLM, from scratch on 23.5T tokens. With approximately 65\% as many pretraining tokens as Qwen3, LLaDA MoE v2 approaches Qwen3 on several knowledge, reasoning, and coding benchmarks. After supervised fine-tuning alone, it outperforms SDAR Chat on seven of eight reasoning and coding benchmarks and remains close to Qwen3 on several tasks. These results establish practical scaling laws and design principles for MoE dLLMs.

0 Citations
0 Influential
4.5 Altmetric
22.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!