2607.08057v1 Jul 09, 2026 cs.LG

효율적인 대규모 언어 모델 서비스 구현을 위한 연구: 시스템 관점에서의 KV 캐시 최적화에 대한 조사

Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization

Peiyu Yang
Peiyu Yang
Citations: 20
h-index: 3
Jiantong Jiang
Jiantong Jiang
Citations: 77
h-index: 5
Rui Zhang
Rui Zhang
Citations: 0
h-index: 0
Feng Liu
Feng Liu
Citations: 1
h-index: 1

대규모 언어 모델(LLM)의 빠른 발전에도 불구하고, LLM 서비스를 제공하는 시스템은 여전히 메모리 사용량이 높고 비용이 많이 듭니다. 자동 회귀 디코딩 과정에서 KV 텐서를 저장하는 키-값(KV) 캐시는 낮은 지연 시간과 높은 처리량을 가진 LLM 추론 서비스 구현에 매우 중요합니다. 본 연구에서는 LLM 서비스를 위한 시스템 관점의 KV 인프라(sKis)에 초점을 맞추어, 기존 연구를 시스템 동작 관점에서 재검토하고, 실행 및 스케줄링(시간적), 배치 및 마이그레이션(공간적), 표현 및 유지(구조적)이라는 세 가지 측면으로 분류합니다. 또한, 상호 작용하는 동작 간의 연관성 및 동작-목표 연결을 분석하여 향후 연구 기회를 제시합니다. 본 연구는 빠르게 발전하고 있는 분야를 체계적으로 정리하여 현대 LLM 서비스 인프라에서 KV 캐시 설계에 대한 이해와 혁신을 위한 기반을 제공합니다.

Original Abstract

Despite the rapid advancements of large language models (LLMs), LLM serving systems remain memory-intensive and costly. The key-value (KV) cache, which stores KV tensors during autoregressive decoding, is crucial for enabling low-latency, high-throughput LLM inference serving. In this survey, we focus on system-aware KV infrastructure for serving LLMs (abbreviated as sKis). We revisit recent work from a system behavior perspective, organizing existing efforts into three dimensions: execution and scheduling (temporal), placement and migration (spatial), and representation and retention (structural). Furthermore, we analyze cross-behavior co-design affinity and behavior-objective links, highlighting future opportunities. Our work systematizes a rapidly evolving area, providing a foundation for understanding and innovating KV cache designs in modern LLM serving infrastructure.

14 Citations
1 Influential
2.5 Altmetric
28.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!