2608.07001v1 Aug 07, 2026 cs.LG

모든 캐시 항목은 그 위치를 가치 있게 한다: KV 캐시 압축을 위한 글로벌 해상도 및 커버리지 할당

Every Cache Entry Earns Its Place: Global Allocation of Resolution and Coverage for KV Cache Compression

Tonghan Wang
Tonghan Wang
Citations: 33
h-index: 4
Haolin Tian
Haolin Tian
Citations: 0
h-index: 0
Yuzhe Liu
Yuzhe Liu
Citations: 0
h-index: 0

대규모 언어 모델(LLM)이 점점 더 긴 문맥을 처리함에 따라, KV 캐시 저장 및 반복적인 접근은 주요 병목 현상이 되고 있습니다. 기존의 KV 캐시 압축 방법은 미리 정의된 고정 규칙에 의존하며, 일반적으로 토큰 제거 또는 병합 방식을 기반으로 개발됩니다. 그 결과, 캐시 리소스는 레이어, 헤드 및 컨텍스트 슬롯 간에 자유롭게 흐르지 못하고, 로컬 해상도와 정보 커버리지를 균형 있게 할당할 수 없습니다. 따라서 우리는 KV 캐시 압축에서 해상도와 커버리지 할당을 위한 글로벌 접근 방식인 GraceKV를 제안하며, 압축 과정을 고정된 캐시 예산 하에서의 글로벌 리소스 할당 문제로 공식화합니다. GraceKV는 각 레이어-KV 헤드-슬롯 조합을 하나의 원자 단위로 취급하고 프로토타입 트리를 구축합니다. 리프 노드는 토큰 수준의 KV 항목에 해당하며, 각 내부 노드는 자식 노드가 커버하는 KV 공간을 단일 프로토타입으로 압축하여 표현합니다. 트리의 서로 겹치지 않는 노드 집합은 하나의 원자 단위를 나타냅니다. 새로운 트리의 루트를 추가하면 정보 커버리지가 확장되고, 선택된 노드를 분할하면 로컬 해상도가 향상됩니다. 모든 후보 액션은 공유 캐시 예산을 놓고 글로벌 경쟁을 합니다. 마지막으로, 모든 트리에서 유지되는 노드는 압축된 KV 캐시를 형성합니다. 이 과정은 전체적으로 원자 단위 간의 캐시 리소스 할당과 해상도와 커버리지 사이의 균형을 적응적으로 결정합니다. GraceKV는 추가적인 학습이 필요 없으며, 전체 압축 및 추론 프로세스는 GPU에서 수행됩니다. 다양한 긴 문맥 작업 및 압축 비율에 대한 체계적인 실험 결과, GraceKV는 32가지 설정 중 24가지에서 가장 우수한 성능을 보였으며, 최대 128배의 압축에도 안정성을 유지했습니다. 이러한 결과는 정보 커버리지와 로컬 해상도를 조율하는 데 있어 글로벌 예산 할당의 효과를 입증합니다.

Original Abstract

As large language models (LLMs) process increasingly long contexts, KV cache storage and repeated access have become a major bottleneck. Existing KV cache compression methods rely on predefined, fixed compression rules and are typically developed around either token eviction or merging. As a result, cache resources can neither flow freely across layers, heads, and context slots, nor be jointly allocated to balance local resolution and information coverage. Therefore, we propose GraceKV, a global approach for the allocation of resolution and coverage in KV cache compression, and formulate the compression process as a global resource allocation problem under a fixed cache budget. GraceKV treats each layer-KV head-slot combination as an atomic unit and builds a prototype tree. Leaf nodes correspond to token-level KV entries, while each internal node uses a single prototype to compress the KV space covered by its children. A set of non-overlapping nodes in the tree forms the representation of an atomic unit. Adding the root of a new tree expands information coverage, whereas splitting a selected node improves local resolution. All candidate actions compete globally for a shared cache budget. Finally, the nodes retained across all trees form the compressed KV cache. This process adaptively determines the allocation of cache resources among atomic units globally and the balance between resolution and coverage. GraceKV requires no additional training, and the entire compression and inference process is performed on the GPU. Systematic experiments across diverse long-context tasks and compression ratios show that GraceKV ranks first in 24 of 32 settings and remains robust up to 128-fold compression. These results validate the effectiveness of global budget allocation in coordinating information coverage and local resolution.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!