2608.03893v1 Aug 04, 2026 cs.LG

LLM 패밀리 간의 교차 모델 KV 캐시 전송: 프리필 재사용을 위한 폐쇄형 선형 매핑

Cross-Model KV Cache Transfer in LLM Families: A Closed-Form Linear Mapping for Prefill Reuse

Rasoul Shafipour
Rasoul Shafipour
Citations: 703
h-index: 13
Maximilian Golub
Maximilian Golub
Citations: 349
h-index: 4
Ritchie Zhao
Ritchie Zhao
Citations: 228
h-index: 5
Ritika Borkar
Ritika Borkar
Citations: 165
h-index: 5
B. Rouhani
B. Rouhani
Citations: 2,705
h-index: 22
Makesh Chandran
Makesh Chandran
Citations: 22
h-index: 3
Mohammad Mahdi Kamani
Mohammad Mahdi Kamani
Pennsylvania State University
Citations: 1,696
h-index: 13
Taekyung Heo
Taekyung Heo
Citations: 362
h-index: 8
Pantea Zardoshti
Pantea Zardoshti
Citations: 615
h-index: 8

실제 환경에서는 비용-성능 균형, 대화 중 전환 및 라우팅을 위해 다양한 크기의 모델이 사용되며, 이러한 전환은 수신 측에서 프리필 과정을 처음부터 다시 수행하도록 요구합니다. 본 연구에서는 교차 모델 KV 캐시 전송 기법을 제안하며, 이를 통해 수신 측에서 소스 모델의 KV 캐시를 재사용하여 프리필 단계를 생략할 수 있습니다. 분석 결과, 일관된 KV 헤드 개수와 헤드당 차원을 갖는 쌍 모델 간의 KV 값에 상당한 선형 구조가 존재하는 것을 확인했습니다. Qwen3 14B->32B 모델에서 하나의 소스 레이어가 타겟 모델의 키 값 변동의 최대 56% (키) 및 32% (값)를 설명하며, 여러 소스 레이어를 사용할 경우 이 비율은 각각 79%와 65%로 증가합니다. 이러한 점을 바탕으로, 각 헤드별로 작동하는 폐쇄형 리지 매퍼를 설계했습니다. 이 매퍼는 세 단계로 구성됩니다. 첫째, 각 타겟 레이어에 대해 가장 예측력이 높은 상위 k개의 소스 레이어를 선택하고, 해당 레이어들의 KV 값을 연결하여 입력으로 사용합니다. 둘째, 매핑 전에 키 값에서 RoPE(Rotary Positional Embedding)를 제거하여 위치 정보에 독립적이고 다양한 컨텍스트 길이에 재사용할 수 있도록 합니다. 셋째, FineWeb-Edu 데이터셋의 1,024 토큰 길이 시퀀스 500개를 사용하여 리지 회귀 모델을 학습합니다. 놀랍게도, 세 개의 패밀리에 속하는 여섯 쌍의 모델에서 이 선형 매퍼는 네 쌍의 경우 수신 모델의 독립적인 프리필 정확도의 73-98%를 유지하며, 나머지 두 쌍에서는 성능이 저하되었습니다. 비선형 MLP(Multi-Layer Perceptron)을 사용하면 실패한 경우에도 HellaSwag 데이터셋에 대한 정확도를 최대 +37 pp만큼 향상시킬 수 있습니다. 제안된 매퍼는 기존 프리필 방식보다 2.7~25배 빠르며, 다중 턴 대화에서도 안정적인 성능을 보여주므로 교차 모델 KV 캐시 전송이 실용적임을 입증합니다.

Original Abstract

Production deployments often swap between different-sized models in a family for cost-quality cascading, mid-conversation switching, and routing, and each swap forces the receiver to repay the prefill from scratch. We propose cross-model KV cache transfer, where the receiver reuses the source's KV cache, skipping prefill. We find that cross-model KV has substantial linear structure across matched-KV pairs, where source and target share KV head count and per-head dimension. On Qwen3 14B->32B, one source layer explains 56% of variance in the target's keys and 32% in values, rising to 79% and 65% with multiple source layers. Building on this, we design a closed-form ridge mapper that operates per head and proceeds in three steps. First, for each target layer we select the top-k most predictive source layers and concatenate their KV as input. Second, we strip RoPE from the keys before mapping, so the fit is position-free and reusable across context lengths. Third, we fit ridge regression on a small calibration set of 500 FineWeb-Edu sequences of 1,024 tokens each. Surprisingly, across six pairs in three families, this linear mapper retains 73-98% of the receiver's standalone-prefill accuracy on four pairs, while two degrade sharply. A nonlinear MLP recovers up to +37 pp HellaSwag retention on the failures. The mapper runs 2.7-25x faster than re-prefill and remains stable across multi-turn handoff, making cross-model KV cache transfer practical.

0 Citations
0 Influential
11 Altmetric
55.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!