2607.02010v1 Jul 02, 2026 cs.AI

InduceKV: KV 메모리를 활용한 멀티모달 LLM의 고정 크기 지속적 적응

InduceKV: Fixed-Footprint Continual Adaptation of Multimodal LLMs via Inducing KV Memories

Canran Xiao
Canran Xiao
Citations: 58
h-index: 4
Qianyu Chen
Qianyu Chen
Citations: 11
h-index: 2
Runxuan Tang
Runxuan Tang
Citations: 1
h-index: 1
Ziteng Feng
Ziteng Feng
Citations: 0
h-index: 0

멀티모달 대규모 언어 모델은 변화하는 작업과 도메인에 적응해야 하지만, 제한된 배포 공간 내에서 지속적인 개선은 반복적인 파라미터 업데이트 또는 증가하는 재생 데이터 저장소로 인해 시간이 지남에 따라 적응 상태가 누적되기 때문에 어렵습니다. 본 연구에서는 고정 크기 지속적 적응을 다룹니다. 즉, 배포되는 적응 상태는 고정된 메모리 예산 내에 유지되며, 기본 모델은 변경되지 않고 작업별 업데이트는 외부적으로 처리됩니다. 우리는 InduceKV라는 검색 기반 방법을 제안합니다. 이 방법은 선택된 각 훈련 데이터 프레임워크를 어텐션 준비 메모리 항목으로 저장하며, 여기에는 고정된 검색 키와 함께 모델의 자기-어텐션 캐시에 추가될 수 있는 간결한 레이어별 키-값(KV) 페이로드가 포함됩니다. 엄격한 메모리 예산 하에서 InduceKV는 이분 수준 선택을 통해 컴팩트한 유도 집합을 구성합니다. 경량 교정 단계를 사용하여 검색을 최적화하고, 선택된 메모리는 현재 작업의 가능성, 앵커 기반 유지 및 고정된 검색 공간에서의 커버리지를 균형 있게 고려합니다. 작업 증분 지시 조정, 지속적인 VQA(Visual Question Answering), 도메인 증분 적응, 그리고 수명 주기 멀티모달 지시 조정 실험에서 InduceKV는 일관적으로 PEFT, MoE, 재생 데이터 저장소 및 프롬프트 검색 기반 방법보다 우수한 성능을 보입니다. 또한, 기본 모델 성능과의 비교, 1단계 CoIN(Contrastive Instruction Tuning), 컴퓨팅 자원 비교, 그리고 확장성 진단을 통해 얻은 결과는 더 강력한 기본 모델, 재생 데이터 저장소만 사용하거나 무제한의 후보 풀이 아닌 경우에도 성능 향상이 가능하다는 것을 보여줍니다.

Original Abstract

Multimodal large language models must adapt to evolving tasks and domains, yet continual improvement under bounded deployment footprint remains difficult because repeated parameter updates or growing replay stores can accumulate adaptation state over time. We study fixed-footprint continual adaptation: the deployed adaptation state is kept under a fixed memory budget, while the backbone model is left unchanged and task-specific updates are externalized. We propose InduceKV, a retrieval-based method that stores each selected training prefix as an attention-ready memory entry, consisting of a frozen retrieval key and compact layerwise key--value (KV) payloads that can be appended to the model's self-attention cache. Under a strict memory budget, InduceKV constructs a compact inducing set through bilevel selection: a lightweight calibration is fit for retrieval, while the selected memory balances current-task likelihood, anchor-based retention, and coverage in the frozen retrieval space. Across task-incremental instruction tuning, continual VQA, domain-incremental adaptation, and lifelong multimodal instruction tuning, InduceKV consistently improves over PEFT, MoE, replay, and prompt-retrieval baselines under matched memory budgets. We further report backbone-matched, stage-1 CoIN, compute-matched, and scalability diagnostics, showing that the gains are not due to a stronger backbone, replay alone, or an unbounded candidate pool.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!