2606.09508v1 Jun 08, 2026 cs.AI

고정에서 동적으로: 엔트로피 기반 적응적 추론을 통한 장문 컨텍스트 LLM

From Rigid to Dynamic: Entropy-Guided Adaptive Inference for Long-Context LLMs

Qing Li
Qing Li
Citations: 151
h-index: 2
Haoyang Li
Haoyang Li
Citations: 451
h-index: 10
Fei Teng
Fei Teng
The Hong Kong University of Science and Technology
Citations: 35
h-index: 3
Zhanchao Xu
Zhanchao Xu
Citations: 149
h-index: 2
Q. Xiao
Q. Xiao
Citations: 36
h-index: 3
Chen Jason Zhang
Chen Jason Zhang
Citations: 48
h-index: 2
Lei Chen
Lei Chen
Citations: 21
h-index: 3

장문 컨텍스트 LLM 추론에 사용되는 기존의 희소 어텐션 및 KV 캐시 압축 방법은 일반적으로 모든 어텐션 헤드에 대해 고정된 희소 패턴 또는 균일한 예산을 적용하여, 헤드와 컨텍스트 간의 상당한 차이를 간과합니다. 우리는 어텐션 헤드에서 두 가지 뚜렷한 엔트로피 패턴을 관찰했습니다. '고정 헤드(Rigid Heads)'는 입력 세그먼트에 걸쳐 엔트로피가 거의 0에 머무르는 반면, '동적 헤드(Dynamic Heads)'는 엔트로피가 크게 변동합니다. 중요한 점은 이러한 유형의 분포가 컨텍스트에 따라 달라지며 사전에 결정할 수 없다는 것입니다. 따라서 우리는 어텐션 엔트로피를 사용하여 개별 헤드 및 세그먼트 수준에서 컴퓨팅 자원을 적응적으로 할당하는 훈련이 필요 없는 프레임워크인 EntropyInfer를 제안합니다. 디코딩 과정에서는 생성된 출력 토큰을 활용하여 가장 중요한 캐시 항목을 식별하고 유지하는 잠재적인 KV 캐시 압축 방식을 도입합니다. Llama, Qwen 및 openPangu 모델 시리즈에 대한 광범위한 실험 결과, EntropyInfer는 SnapKV, AdaKV 및 CritiPrefill과 같은 기존 방법보다 일관되게 성능이 우수하며, 전체 어텐션과 비교하여 품질 저하를 최소화하면서 10만 토큰 이상의 길이에서 최대 2.39배의 엔드투엔드 속도 향상을 달성합니다. 관련 코드는 https://github.com/SHA-4096/EntropyInfer 에서 확인할 수 있습니다.

Original Abstract

Existing sparse attention and KV cache compression methods for long-context LLM inference typically apply fixed sparsity patterns or uniform budgets across all attention heads, overlooking the substantial variation in attention behavior among heads and contexts. We observe two distinct entropy patterns among attention heads: Rigid Heads, whose entropy stays near zero across input segments, and Dynamic Heads, whose entropy fluctuates significantly. Crucially, the distribution of these types is context-dependent and cannot be predetermined offline. We therefore propose EntropyInfer, a training-free framework that uses attention entropy to adaptively allocate compute at the granularity of individual heads and segments during prefilling. For decoding, we introduce a latent KV cache compression scheme that leverages generated output tokens, rather than prefill tokens alone, to identify and retain the most critical cache entries. Extensive experiments on Llama, Qwen and openPangu model series show that EntropyInfer consistently outperforms baselines including SnapKV, AdaKV, and CritiPrefill, achieving up to 2.39$\times$ end-to-end speedup beyond 100k tokens with minimal quality degradation compared to full attention. The code is released in https://github.com/SHA-4096/EntropyInfer.

0 Citations
0 Influential
25 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!