LLM 패밀리 간의 교차 모델 KV 캐시 전송: 프리필 재사용을 위한 폐쇄형 선형 매핑
Cross-Model KV Cache Transfer in LLM Families: A Closed-Form Linear Mapping for Prefill Reuse
실제 환경에서는 비용-성능 균형, 대화 중 전환 및 라우팅을 위해 다양한 크기의 모델이 사용되며, 이러한 전환은 수신 측에서 프리필 과정을 처음부터 다시 수행하도록 요구합니다. 본 연구에서는 교차 모델 KV 캐시 전송 기법을 제안하며, 이를 통해 수신 측에서 소스 모델의 KV 캐시를 재사용하여 프리필 단계를 생략할 수 있습니다. 분석 결과, 일관된 KV 헤드 개수와 헤드당 차원을 갖는 쌍 모델 간의 KV 값에 상당한 선형 구조가 존재하는 것을 확인했습니다. Qwen3 14B->32B 모델에서 하나의 소스 레이어가 타겟 모델의 키 값 변동의 최대 56% (키) 및 32% (값)를 설명하며, 여러 소스 레이어를 사용할 경우 이 비율은 각각 79%와 65%로 증가합니다. 이러한 점을 바탕으로, 각 헤드별로 작동하는 폐쇄형 리지 매퍼를 설계했습니다. 이 매퍼는 세 단계로 구성됩니다. 첫째, 각 타겟 레이어에 대해 가장 예측력이 높은 상위 k개의 소스 레이어를 선택하고, 해당 레이어들의 KV 값을 연결하여 입력으로 사용합니다. 둘째, 매핑 전에 키 값에서 RoPE(Rotary Positional Embedding)를 제거하여 위치 정보에 독립적이고 다양한 컨텍스트 길이에 재사용할 수 있도록 합니다. 셋째, FineWeb-Edu 데이터셋의 1,024 토큰 길이 시퀀스 500개를 사용하여 리지 회귀 모델을 학습합니다. 놀랍게도, 세 개의 패밀리에 속하는 여섯 쌍의 모델에서 이 선형 매퍼는 네 쌍의 경우 수신 모델의 독립적인 프리필 정확도의 73-98%를 유지하며, 나머지 두 쌍에서는 성능이 저하되었습니다. 비선형 MLP(Multi-Layer Perceptron)을 사용하면 실패한 경우에도 HellaSwag 데이터셋에 대한 정확도를 최대 +37 pp만큼 향상시킬 수 있습니다. 제안된 매퍼는 기존 프리필 방식보다 2.7~25배 빠르며, 다중 턴 대화에서도 안정적인 성능을 보여주므로 교차 모델 KV 캐시 전송이 실용적임을 입증합니다.
Production deployments often swap between different-sized models in a family for cost-quality cascading, mid-conversation switching, and routing, and each swap forces the receiver to repay the prefill from scratch. We propose cross-model KV cache transfer, where the receiver reuses the source's KV cache, skipping prefill. We find that cross-model KV has substantial linear structure across matched-KV pairs, where source and target share KV head count and per-head dimension. On Qwen3 14B->32B, one source layer explains 56% of variance in the target's keys and 32% in values, rising to 79% and 65% with multiple source layers. Building on this, we design a closed-form ridge mapper that operates per head and proceeds in three steps. First, for each target layer we select the top-k most predictive source layers and concatenate their KV as input. Second, we strip RoPE from the keys before mapping, so the fit is position-free and reusable across context lengths. Third, we fit ridge regression on a small calibration set of 500 FineWeb-Edu sequences of 1,024 tokens each. Surprisingly, across six pairs in three families, this linear mapper retains 73-98% of the receiver's standalone-prefill accuracy on four pairs, while two degrade sharply. A nonlinear MLP recovers up to +37 pp HellaSwag retention on the failures. The mapper runs 2.7-25x faster than re-prefill and remains stable across multi-turn handoff, making cross-model KV cache transfer practical.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.