UBEP: 프로덕션 슈퍼파드를 위한 전문가 병렬 통신 라이브러리의 재설계
UBEP: Re-architecting Expert Parallelism Communication Library for Production Superpods
NVIDIA의 NVL72/576 및 Huawei의 CloudMatrix384와 같은 고대역폭 슈퍼파드에서 Mixture-of-Experts (MoE) 모델을 사용하는 것은 단순히 인터커넥트 대역폭 이상의 중요한 과제를 야기합니다. 이러한 시스템은 통합된 글로벌 주소 공간과 높은 대역폭 구조를 제공하지만, 희소 MoE 통신의 잠재력을 완전히 활용하는 데에는 세 가지 근본적인 병목 현상이 존재합니다: (1) 상호 의존적인 통신 단계를 일괄 동기화 병렬 (BSP) 방식으로 조율할 때 발생하는 엄격한 실행 직렬화; (2) 높은 인터커넥트 대역폭에 비례하여 확장되지 않는 과도한 동기화 오버헤드; 그리고 (3) 거리 무시적인 스케줄링으로 인해 발생하는 심각한 부하 불균형으로 인한 불규칙한 토큰 트래픽. 이러한 병목 현상을 해결하기 위해, 우리는 UBEP (Unified-Bus Expert Parallelism)이라는 프로덕션 환경에 적합한 통신 라이브러리를 소개합니다. 이 라이브러리는 최신 슈퍼파드 아키텍처를 위한 MoE의 All-to-All 연산을 재설계했습니다. 대규모 실험을 통해 UBEP는 All-to-All 지연 시간을 최대 52.4%까지, 그리고 MoE 추론 시 출력 토큰당 시간 (TPOT)을 최대 11.1%까지 줄였습니다.
The deployment of Mixture-of-Experts (MoE) models on production high-bandwidth superpods, such as NVIDIA's NVL72/576 and Huawei's CloudMatrix384, introduces critical challenges beyond raw interconnect bandwidth. While these systems provide unified global address spaces and high-bandwidth fabrics, their full potential for sparse MoE communication is hindered by three fundamental bottlenecks: (1) Strict execution serialization imposed by coarse-grained Bulk Synchronous Parallel (BSP) orchestration of interdependent communication phases; (2) Prohibitive synchronization overhead that fails to scale alongside high interconnect bandwidth; and (3) Severe load imbalance resulting from distance-agnostic scheduling of irregular token traffic. To eliminate these bottlenecks, we introduce UBEP (Unified-Bus Expert Parallelism), a production-ready communication library that rethinks MoE's All-to-All primitives for modern superpod architectures. Through large scale experiments, UBEP reduces All-to-All latency by up to 52.4% and MoE inference Time Per Output Token (TPOT) by up to 11.1%.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.