2608.02989v1 Aug 04, 2026 cs.LG

AcceptMoE: 효율적인 MoE 추론 디코딩을 위한 약속 가중 자가 크기 조정 검증 전문가 집합

AcceptMoE: Commitment-Weighted Self-Sizing Verifier Expert Sets for Efficient MoE Speculative Decoding

Shuang Liang
Shuang Liang
Citations: 11
h-index: 2
H. Chen
H. Chen
Citations: 107
h-index: 6
Zhiwen Mo
Zhiwen Mo
Citations: 18
h-index: 3
Lingxiao Ma
Lingxiao Ma
Citations: 1,115
h-index: 13
Wayne Luk
Wayne Luk
Citations: 4
h-index: 2
Qianzhou Wang
Qianzhou Wang
Citations: 1
h-index: 1
Guoyu Li
Guoyu Li
Citations: 1
h-index: 1

추론 디코딩은 목표 모델의 단일 순방향 연산을 통해 여러 후보 토큰 트리를 검증합니다. 그러나 혼합 전문가(MoE) 기반의 목표 모델에서, 병렬 검증은 모든 트리 노드에 의해 선택된 전문가들의 합집합을 활성화할 수 있습니다. 하지만 실제로 허용되는 출력에 도달하는 노드는 일부분일 뿐입니다. 따라서 토큰 수, 활성화된 전문가 집합 크기, 그리고 전문가 가중치 트래픽은 서로 다른 비용 지표이며, 토큰 작업량을 줄이는 것이 반드시 전문가 집합 크기를 비례적으로 감소시키지 않을 수 있습니다. 또한 오프로딩 시에는 전송 트래픽도 캐시 잔존 여부에 따라 달라집니다. 본 논문에서는 AcceptMoE를 제안합니다. AcceptMoE는 검증 측면에서 작동하는 전문가 선택기로, 목표 라우터 점수를 오프라인으로 추정한 약속 확률과 결합하여 각 검증 블록에 대해 사용할 수 있는 전문가 수를 자동으로 조정합니다. 이를 통해 사용자가 지정한 전문가 예산의 필요성을 없앱니다. 오프로딩 환경에서는 AcceptMoE가 자연 경로를 예측하고 해당 전문가 가중치를 미리 가져오는 대신, 캐시 잔존 여부에 따라 전문가 자격을 결정합니다. 목표-전문가 자격 요건을 제한하면 모델 분포가 변경될 수 있지만, 3가지 MoE 목표 및 4개의 벤치마크를 포함하는 12개의 모델-태스크 쌍에서 AcceptMoE의 평균 정확도는 자연 라우팅을 사용하는 EAGLE-3 추론 디코딩보다 0.27%p 낮은 것으로 나타났습니다. SGLang과 배치 크기 1로 실행했을 때, AcceptMoE는 모든 전문가 가중치를 GPU 메모리에 저장한 기준 모델보다 처리량을 1.29배 향상시켰으며, 물리적 전문가 오프로딩 환경에서는 2.06배의 성능을 보였습니다. 또한 호스트-장치 트래픽을 73.6%에서 77.1%까지 줄였습니다.

Original Abstract

Speculative decoding verifies a tree of draft tokens in one target-model forward pass. For a mixture-of-experts (MoE) target, however, parallel verification can activate the union of the experts selected by all tree nodes, even though only a small subset of those nodes reaches the accepted output. Token count, activated-expert union size, and expert-weight traffic are therefore distinct cost measures: reducing the token workload need not shrink the expert union proportionally, and under offloading, transfer traffic also depends on cache residency. We introduce AcceptMoE, a verifier-side expert selector that combines target-router scores with offline-estimated commitment probabilities and automatically adjusts the number of eligible experts for each verification block, eliminating the need for a user-specified expert budget. Under offloading, AcceptMoE conditions expert eligibility on cache residency instead of predicting natural routes and prefetching the corresponding expert weights. Although constraining target-expert eligibility changes the model distribution, across 12 model-task pairs spanning three MoE targets and four benchmarks, AcceptMoE's mean accuracy is 0.27 percentage points lower than that of EAGLE-3 speculative decoding with natural routing. Served with SGLang at batch size one, it reaches 1.290 times the throughput of this baseline with all expert weights in GPU memory, and 2.06 times under physical expert offloading, while reducing host-to-device traffic by 73.6 percent to 77.1 percent.

0 Citations
0 Influential
6.5 Altmetric
32.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!