2608.02901v1 Aug 03, 2026 cs.LG

AnchorKV: 앵커 잔차 키-값 캐시 압축

AnchorKV: Anchor-Residual KV Cache Compression

Yuval Sieradzki
Yuval Sieradzki
Citations: 19
h-index: 2
M. Khalaf
M. Khalaf
Citations: 561
h-index: 10
Yara Shamshoum
Yara Shamshoum
Citations: 15
h-index: 1
Nitzan Hodos
Nitzan Hodos
Citations: 18
h-index: 2
Assaf Schuster
Assaf Schuster
Citations: 20
h-index: 2

키-값(KV) 캐시는 긴 문맥을 처리하는 대규모 언어 모델(LLM) 추론 과정에서 주요 메모리 병목 현상을 야기합니다. 기존 방법들은 이 문제를 해결하기 위해 서로 다른 접근 방식을 사용합니다. 삭제 방식은 토큰을 영구적으로 제거하여 성능 저하를 일으키는 반면, 양자화 방식은 모든 토큰을 낮은 정밀도로 유지하지만 압축률이 제한적입니다. 본 논문에서는 AnchorKV라는 새로운 압축 기법을 제안합니다. AnchorKV는 단 하나의 토큰도 삭제하지 않고 캐시 크기를 20배까지 줄일 수 있습니다. AnchorKV는 정확하게 저장된 소량의 '앵커'를 사용하여 캐시를 표현하고, 나머지 모든 토큰은 가장 유사한 앵커를 통해 표현하며, 모델 출력에 가장 큰 영향을 미치는 토큰만 추가적으로 정제합니다. AnchorKV는 다양한 모델과 데이터셋에서 일관된 정확도를 유지하며, 70B 규모의 모델에서도 전체 캐시 성능의 99%를 보존하면서도 전체 문맥을 기존 비용의 일부로 유지할 수 있습니다.

Original Abstract

The key-value (KV) cache is the primary memory bottleneck in long-context LLM inference. Existing approaches attack it from opposite ends: eviction methods permanently discard tokens, degrading performance whenever a discarded token later proves essential, while quantization methods retain all tokens at low precision but offer limited compression. We propose AnchorKV, a compression scheme that shrinks the cache by $20\times$ without discarding a single token. AnchorKV represents the cache using a small set of anchors stored exactly, expresses every other token through its most similar anchor, and refines only those whose approximation most affects the model's output. AnchorKV consistently preserves accuracy across models and datasets, retaining 99% of the full-cache score at the 70B scale, while keeping the entire context at a fraction of its cost.

0 Citations
0 Influential
5 Altmetric
25.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!