2603.01058v1 Mar 01, 2026 cs.AR

TriMoE: AMX 지원 CPU 및 DIMM-NDP를 활용한 고처리량 MoE 추론을 위한 GPU 증강 및 오프로딩

TriMoE: Augmenting GPU with AMX-Enabled CPU and DIMM-NDP for High-Throughput MoE Inference via Offloading

Yudong Pan
Yudong Pan
Citations: 92
h-index: 3
Mengdi Wang
Mengdi Wang
Citations: 224
h-index: 8
Yintao He
Yintao He
Citations: 172
h-index: 8
Tian Han
Tian Han
Citations: 55
h-index: 5
Lian Liu
Lian Liu
Citations: 50
h-index: 4
Shixin Zhao
Shixin Zhao
Citations: 25
h-index: 2
Zhirong Chen
Zhirong Chen
Citations: 115
h-index: 4
Cangyuan Li
Cangyuan Li
Citations: 126
h-index: 5
Yinhe Han
Yinhe Han
Citations: 59
h-index: 4
Ying Wang
Ying Wang
Citations: 8
h-index: 1

대규모 Mixture-of-Experts (MoE) 모델을 비용 효율적으로 배포하기 위해서는 오프로딩 기반의 단일 GPU 이기종 추론이 필수적입니다. GPU-CPU 아키텍처에서 콜드(cold) 전문가를 오프로딩할 때 호스트 메모리 대역폭의 제약이 있는 반면, 새로운 GPU-NDP 아키텍처는 DIMM-NDP를 활용하여 핫(hot)하지 않은 전문가를 오프로딩합니다. 그러나 핫하지 않은 전문가들은 균일한 메모리 바운드 그룹이 아니며, 상당수의 웜(warm) 전문가들이 높은 GPU I/O 지연으로 인해 심각한 성능 저하를 겪으면서도 NDP의 연산 처리량을 완전히 활용하지 못하는 중요한 성능 격차가 존재합니다. 본 논문에서는 AMX 지원 CPU를 활용하여 핫, 웜, 콜드 전문가를 최적의 연산 장치에 정확하게 매핑하여 이러한 격차를 해소하는 새로운 GPU-CPU-NDP 아키텍처인 TriMoE를 제안합니다. 또한, 병목 현상을 고려한 전문가 스케줄링 정책과 예측 기반의 동적 레이아웃/재분배 방식을 도입했습니다. 실험 결과, TriMoE는 최첨단 솔루션 대비 최대 2.83배의 성능 향상을 달성했습니다.

Original Abstract

To deploy large Mixture-of-Experts (MoE) models cost-effectively, offloading-based single-GPU heterogeneous inference is crucial. While GPU-CPU architectures that offload cold experts are constrained by host memory bandwidth, emerging GPU-NDP architectures utilize DIMM-NDP to offload non-hot experts. However, non-hot experts are not a homogeneous memory-bound group: a significant subset of warm experts exists is severely penalized by high GPU I/O latency yet can saturate NDP compute throughput, exposing a critical compute gap. We present TriMoE, a novel GPU-CPU-NDP architecture that fills this gap by synergistically leveraging AMX-enabled CPU to precisely map hot, warm, and cold experts onto their optimal compute units. We further introduce a bottleneck-aware expert scheduling policy and a prediction-driven dynamic relayout/rebalancing scheme. Experiments demonstrate that TriMoE achieves up to 2.83x speedup over state-of-the-art solutions.

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!