2605.28207v1 May 27, 2026 cs.CL

Mixture-of-Experts 모델을 Dense Language Model로 변환하는 기술: 가지치기 및 지식 증류

Pruning and Distilling Mixture-of-Experts into Dense Language Models

Gyeongman Kim
Gyeongman Kim
Citations: 180
h-index: 5
Jihun Yun
Jihun Yun
Citations: 15
h-index: 3
Haechan Kim
Haechan Kim
Citations: 8
h-index: 2
Junhyuck Kim
Junhyuck Kim
Citations: 54
h-index: 3
J. Bae
J. Bae
Citations: 0
h-index: 0
Jaewoong Cho
Jaewoong Cho
Citations: 49
h-index: 3

Mixture-of-Experts (MoE)는 최첨단 언어 모델의 주요 아키텍처이지만, 모든 전문가 파라미터를 메모리에 로드해야 하므로 메모리 제약이 있는 환경에서의 배포에는 적합하지 않습니다. 기존의 압축 방법은 전문가 수를 줄이지만, 여전히 동일한 근본적인 한계를 가진 MoE 모델을 유지합니다. 본 연구에서는 학습된 MoE 모델을 표준 Dense 아키텍처로 변환하는 최초의 체계적인 프레임워크를 제시합니다. 구체적으로, 전문가들을 점수화하고 선택 및 그룹화하여 Dense Feed Forward Network (FFN)으로 연결한 다음, MoE 모델로부터 지식 증류를 통해 성능을 개선합니다. Qwen3-30B-A3B 모델에서 선택된 전문가 수에 따라 7가지 점수화 방법, 5가지 그룹화 방법 및 2가지 크기 조정 방법을 사용하여 총 350개의 구성을 평가했습니다. 그 결과, 점수화 방법의 선택이 가장 큰 영향을 미치는 것으로 나타났으며, 제안하는 다양성 기반 점수화 방법은 Qwen3-30B-A3B, DeepSeek-V2-Lite 및 GPT-OSS-20B 모델에서 기존 방법보다 우수한 성능을 보였습니다. 동일한 파라미터 수로 비교했을 때, MoE를 Dense 모델로 변환하는 방식이 ~40억 토큰으로 지식 증류를 수행하고 1.6배 빠른 학습 속도를 보이는 Dense-to-Dense 가지치기 방법에 비해 평균 다운스트림 정확도에서 +6.3%p 더 높은 성능을 보였습니다.

Original Abstract

Mixture-of-Experts (MoE) is now the dominant architecture for frontier language models, yet it requires all expert parameters to be loaded in memory, making it less preferable for memory-constrained deployment. Existing compression methods reduce the number of experts but the output remains an MoE model with the same fundamental limitation. We present the first systematic framework for converting a trained MoE into a standard fully dense architecture: experts are scored, selected, and grouped, then concatenated into a dense FFN and refined by knowledge distillation from the MoE teacher. We evaluate 7 scoring, 5 grouping, and 2 magnitude scaling methods across a range of selected expert counts on Qwen3-30B-A3B, yielding 350 configurations. We find that the choice of scoring method is the most impactful, with our novel diversity-aware scoring consistently outperforming prior methods on Qwen3-30B-A3B, DeepSeek-V2-Lite, and GPT-OSS-20B. Under a controlled comparison at matched parameter count, MoE-to-dense outperforms dense-to-dense pruning by +6.3 pp in average downstream accuracy after ~4B-token distillation at 1.6x faster training wall-clock speed.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!