2605.28302v1 May 27, 2026 cs.LG

얼마나 더 분산화할 수 있을까? 효율적인 MoE LLM 서비스를 위한 어텐션-FFN 분산화 설계 공간 탐색

How Far Can Disaggregation Go? A Design-Space Exploration of Attention-FFN Disaggregation for Efficient MoE LLM Serving

Suvinay Subramanian
Suvinay Subramanian
Citations: 1,944
h-index: 18
Sarbartha Banerjee
Sarbartha Banerjee
Citations: 111
h-index: 6
A. Bambhaniya
A. Bambhaniya
Citations: 140
h-index: 7
Tuhin Khare
Tuhin Khare
Citations: 52
h-index: 3
S. Srinivasan
S. Srinivasan
Citations: 2,352
h-index: 15
Souvik Kundu
Souvik Kundu
Citations: 245
h-index: 8
Midhilesh Elavazhagan
Midhilesh Elavazhagan
Citations: 54
h-index: 4
William Won
William Won
Citations: 279
h-index: 7
Amir Yazdanbakhsh
Amir Yazdanbakhsh
Citations: 3,361
h-index: 7
Tushar Krishna
Tushar Krishna
Citations: 73
h-index: 5
Hanjiang Wu
Hanjiang Wu
Citations: 27
h-index: 3
Madhu Kumar
Madhu Kumar
Citations: 20
h-index: 2

최근의 대규모 언어 모델(LLM) 추론은 모델 크기가 증가하고 TTFT (Time to First Token) 및 TPOT (Throughput per Optimization Target)과 같은 서비스 수준 목표를 충족하기 위해 점진적으로 분산화되어 왔습니다. 초기에는 청크 기반 프리필 방식, 이후 프리필-디코딩(P/D) 분산 방식으로 발전했으며, 가장 최근에는 연산자 레벨의 어텐션-FFN 분산(AFD) 방식을 사용하고 있습니다. 이러한 경향은 특히 메모리 병목 현상이 발생하는 어텐션, 연산 집약적인 전문가 FFN (Feed Forward Network), 그리고 MoE (Mixture of Experts) 배분/결합 통신으로 인해 고유한 리소스 요구 사항을 갖는 MoE 모델에 매우 중요합니다. AFD는 어텐션과 MoE-FFN 실행을 별도의 GPU 그룹에 배치하여 이러한 이질성을 더욱 명확하게 보여줍니다. 분산화의 각 단계는 워크로드 특성, 리소스 할당 및 상호 연결 토폴로지에 걸쳐 설계 공간을 심화시키며, 핵심적인 질문은 '각 단계가 실제로 언제 효과를 발휘하는가?'입니다. 우리는 입력/출력 시퀀스 길이, 프리픽스 KV 재사용, 사용자별 지연 시간 제약 조건을 포함한 다양한 실제 워크로드에 대한 MoE 추론의 이러한 절충점을 체계적으로 분석합니다. 청크 기반 프리필 및 P/D 분산 방식을 기준으로, 온디바이스 커널 측정과 고정밀 네트워크 시뮬레이션을 결합하는 프레임워크를 통해 AFD의 장점과 한계를 대규모로 연구합니다. 엄격한 TTFT/TPOT 서비스 수준 목표(SLO) 하에서, AFD는 DeepSeek-V3.2 모델에서 채팅, 코딩 및 에이전트 코딩 워크로드에 대해 약 4k 토큰/초의 시스템 처리량을 유지하며, 이는 AFD를 사용하지 않는 환경에서는 구현 불가능합니다. 우리는 처리량과 상호 작용성을 동시에 최적화하기 위한 구체적인 지침을 제시하며, 여기에는 워크로드 및 모델 아키텍처에 따라 GPU에서 어텐션과 FFN을 어떻게 분할해야 하는지에 대한 내용이 포함됩니다. 이러한 지침은 현재의 랙 및 클러스터 규모 배포뿐만 아니라 향후 분산 AI 인프라를 위한 설계 원칙을 제공합니다.

Original Abstract

Modern large language model (LLM) inference has progressively disaggregated to keep pace with growing model sizes and tight TTFT and TPOT service-level objectives: from chunked-prefill aggregation, to prefill-decode (P/D) disaggregation, and most recently to operator-level Attention-FFN Disaggregation (AFD). This trend is especially important for mixture-of-experts (MoE) models, where memory-bound attention, compute-intensive expert FFNs, and MoE dispatch/combine communication create distinct resource demands. AFD further exposes this heterogeneity by placing attention and MoE-FFN execution on separate GPU groups. Each level of disaggregation deepens the scheduling design space across workload characteristics, resource allocation, and interconnect topology, raising the central question: when does each level actually pay off? We systematically characterize this trade-off for MoE inference across realistic workloads spanning input/output sequence lengths, prefix-KV reuse, and per-user latency constraints. Using chunked-prefill and P/D disaggregation as baselines, we study the benefits and limits of AFD at scale through a framework that fuses on-device kernel measurements with high-fidelity network simulation. Under strict TTFT/TPOT SLOs, AFD sustains around 4k tokens/s of system throughput on DeepSeek-V3.2 across chat, coding, and agentic-coding workloads, where non-AFD deployments are infeasible. We distill concrete takeaways for jointly optimizing throughput and interactivity, including how to partition attention and FFN across GPUs as a function of workload and model architecture, providing design principles for current rack- and cluster-scale deployments as well as future disaggregated AI infrastructure.

2 Citations
0 Influential
9 Altmetric
47.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!