2605.27358v1 May 26, 2026 cs.LG

MobileMoE: 온디바이스 혼합 전문가(Mixture of Experts) 모델의 확장

MobileMoE: Scaling On-Device Mixture of Experts

Ernie Chang
Ernie Chang
Citations: 511
h-index: 6
Zechun Liu
Zechun Liu
Citations: 2,171
h-index: 10
Raghuraman Krishnamoorthi
Raghuraman Krishnamoorthi
Citations: 5,227
h-index: 18
Yanbei Chen
Yanbei Chen
Citations: 333
h-index: 9
Han Huang
Han Huang
Citations: 3
h-index: 1
Jacob Szwejbka
Jacob Szwejbka
Citations: 0
h-index: 0
Digant Desai
Digant Desai
Citations: 31
h-index: 1
Vikas Chandra
Vikas Chandra
Citations: 1,463
h-index: 11

혼합 전문가(MoE)는 수백억 개의 파라미터를 가진 언어 모델의 사실상 표준 아키텍처로 자리 잡았지만, 10억 개 미만의 규모에서 온디바이스 배포에 있어 MoE의 장점은 아직 충분히 탐구되지 않았습니다. 이러한 간극을 해소하기 위해, 저희는 MobileMoE를 제안합니다. MobileMoE는 활성 파라미터 수가 수십억 개(0.3-0.9B 활성 및 1.3-5.3B 총 파라미터)인 온디바이스 MoE 언어 모델 패밀리로, 온디바이스 LLM에 대한 새로운 성능 지평을 제시합니다. 먼저, 저희는 모바일 메모리와 컴퓨팅 제약 조건 하에서 MoE 아키텍처를 공동으로 최적화하는 온디바이스 MoE 확장 법칙을 정립했습니다. 이를 통해 메모리와 컴퓨팅 모두에서 최적인 '온디바이스 이상적인 지점'을 발견했는데, 이는 적절한 희소성과 미세 조정된 공유 전문가를 특징으로 합니다. 도출된 아키텍처를 기반으로, 저희는 사전 훈련, 중간 훈련, 명령어 파인튜닝 및 양자화 인식 훈련의 네 단계로 구성된 방법을 사용하여 오픈 소스 데이터 세트에서 MobileMoE 모델을 학습했습니다. 14개의 벤치마크 테스트 결과, MobileMoE는 선도적인 온디바이스 밀집 LLM과 동등하거나 뛰어난 성능을 보이면서 추론에 필요한 FLOPs가 최대 2-4배 적습니다. 또한, 최고 수준의 MoE 모델인 OLMoE-1B-7B와 비교하여 최대 60% 더 적은 파라미터로 동일하거나 더 나은 성능을 달성했습니다. 모바일 배포를 위한 마지막 단계를 해결하기 위해, 저희는 상용 스마트폰에서 최초의 효율적인 MoE 추론 기능을 제공하며, 온디바이스 프로파일링을 통해 상세한 정보를 제공합니다. INT4 가중치 메모리 측면에서, MobileMoE-S는 밀집 모델인 MobileLLM-Pro에 비해 사전 채우기(prefill) 속도가 1.8배에서 3.8배 빠르고, 디코딩 속도는 2.2배에서 3.4배 빠릅니다.

Original Abstract

Mixture-of-Experts (MoE) has become the de facto architecture for hundred-billion-parameter language models, yet its advantages at sub-billion scales for on-device deployment remain largely unexplored. To close this gap, we present MobileMoE, a family of on-device MoE language models with sub-billion active parameters (0.3-0.9B active and 1.3-5.3B total) that establish a new Pareto frontier for on-device LLMs. We first formulate an on-device MoE scaling law that jointly optimizes MoE architecture under mobile memory and compute constraints, identifying an on-device sweet spot - moderate sparsity with fine-grained and shared experts - that is simultaneously memory and compute-optimal. Building on the derived architectures, we train MobileMoE with a four-stage recipe covering pre-training, mid-training, instruction fine-tuning, and quantization-aware training, all on open-source datasets. Across 14 benchmarks, MobileMoE matches or exceeds leading on-device dense LLMs with 2-4$\times$ fewer inference FLOPs, and matches or surpasses the state-of-the-art MoE OLMoE-1B-7B with up to 60% fewer parameters. To bridge the last mile to mobile deployment, we provide the first efficient MoE inference on commodity smartphones with comprehensive on-device profiling. At comparable INT4 weight memory, MobileMoE-S delivers $1.8$-$3.8\times$ faster prefill and $2.2$-$3.4\times$ faster decode than the dense baseline MobileLLM-Pro.

0 Citations
0 Influential
9 Altmetric
45.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!