2601.20309v1 Jan 28, 2026 cs.DC

SuperInfer: 초급속 칩(Superchip) 기반 LLM 추론을 위한 SLO 인지 로터리 스케줄링 및 메모리 관리

SuperInfer: SLO-Aware Rotary Scheduling and Memory Management for LLM Inference on Superchips

Jiahuan Yu
Jiahuan Yu
Citations: 7
h-index: 2
Mingtao Hu
Mingtao Hu
Citations: 29
h-index: 3
Zichao Lin
Zichao Lin
Citations: 146
h-index: 4
Minjia Zhang
Minjia Zhang
Citations: 75
h-index: 4

대규모 언어 모델(LLM) 서비스는 엄격한 지연 시간 서비스 수준 목표(SLO)와 제한된 GPU 메모리 용량 사이의 근본적인 긴장 관계에 직면합니다. 높은 요청 빈도가 KV 캐시 예산을 초과하면 기존 LLM 추론 시스템은 종종 심각한 헤드-오브-라인(HOL) 블로킹을 겪습니다. 이전 연구에서는 PCIe 기반 오프로딩을 탐구했지만, 이러한 접근 방식은 높은 요청 빈도에서 응답성을 유지할 수 없으며, 종종 엄격한 첫 번째 토큰 시간(TTFT) 및 토큰 간 시간(TBT) SLO를 충족하지 못합니다. 본 논문에서는 NVLink-C2C를 통해 밀접하게 결합된 GPU-CPU 아키텍처를 가진 최신 Superchip(예: NVIDIA GH200)을 위한 고성능 LLM 추론 시스템인 SuperInfer를 소개합니다. SuperInfer는 첫 번째 프로액티브, SLO 인지 로터리 스케줄러인 RotaSched을 도입하여 Superchip에서 응답성을 유지하고, NVLink-C2C를 통한 양방향 전송을 가능하게 하는 최적화된 로테이션 엔진인 DuplexKV를 제공합니다. GH200에서 다양한 모델과 데이터 세트를 사용하여 평가한 결과, SuperInfer는 최첨단 시스템과 비교하여 동등한 TBT 및 처리량을 유지하면서 TTFT SLO 달성률을 최대 74.7% 향상시키는 것으로 나타났습니다. 이는 SLO 인지 스케줄링 및 메모리 공동 설계가 Superchip의 잠재력을 최대한 활용하여 응답성이 뛰어난 LLM 서비스를 제공할 수 있음을 보여줍니다.

Original Abstract

Large Language Model (LLM) serving faces a fundamental tension between stringent latency Service Level Objectives (SLOs) and limited GPU memory capacity. When high request rates exhaust the KV cache budget, existing LLM inference systems often suffer severe head-of-line (HOL) blocking. While prior work explored PCIe-based offloading, these approaches cannot sustain responsiveness under high request rates, often failing to meet tight Time-To-First-Token (TTFT) and Time-Between-Tokens (TBT) SLOs. We present SuperInfer, a high-performance LLM inference system designed for emerging Superchips (e.g., NVIDIA GH200) with tightly coupled GPU-CPU architecture via NVLink-C2C. SuperInfer introduces RotaSched, the first proactive, SLO-aware rotary scheduler that rotates requests to maintain responsiveness on Superchips, and DuplexKV, an optimized rotation engine that enables full-duplex transfer over NVLink-C2C. Evaluations on GH200 using various models and datasets show that SuperInfer improves TTFT SLO attainment rates by up to 74.7% while maintaining comparable TBT and throughput compared to state-of-the-art systems, demonstrating that SLO-aware scheduling and memory co-design unlocks the full potential of Superchips for responsive LLM serving.

2 Citations
0 Influential
2 Altmetric
12.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!