2606.23063v1 Jun 22, 2026 cs.CV

리플레이 없는 지속적 멀티모달 LLM을 위한 어텐션 스펙트럼 정규화

Attention-Spectrum Regularization for Replay-Free Continual Multimodal LLMs

Canran Xiao
Canran Xiao
Citations: 58
h-index: 4
Chuangxin Zhao
Chuangxin Zhao
Citations: 22
h-index: 3
Yanbiao Ma
Yanbiao Ma
Citations: 63
h-index: 4
Yang Liu
Yang Liu
Citations: 54
h-index: 2
Jun Xia
Jun Xia
Citations: 32
h-index: 2
Siyuan Ma
Siyuan Ma
Citations: 4,592
h-index: 13
Mengyao Lyu
Mengyao Lyu
Citations: 259
h-index: 7
Guiguang Ding
Guiguang Ding
Citations: 5,267
h-index: 15

멀티모달 대규모 언어 모델(MLLM)은 시각 도메인, 질문 유형 및 사용자 지침의 비정상적인 스트림에 적응해야 하는 경우가 점점 더 많지만, 지속적인 미세 조정은 종종 이전에 습득한 멀티모달 능력을 심각하게 잊게 만듭니다. 기존의 지속적인 비전-언어 방법은 주로 출력 값을 보존하거나, 과거 데이터를 재생하거나 가짜 데이터를 사용하거나, 임베딩 공간의 기하학적 구조를 정규화하거나, 작업별 매개변수를 할당하지만, 이러한 방법들은 적응 과정에서 이전 능력을 지원하는 내부 교차 모달 어텐션 패턴이 어떻게 변화하는지에 대한 제한적인 제어를 제공합니다. 우리는 어텐션 스펙트럼 정규화(ASR)라는 리플레이가 필요 없는 지속적 학습 프레임워크를 제안합니다. ASR은 교차 모달 어텐션의 기술 기반 구조를 보존합니다. ASR은 교차 어텐션 맵을 2차원 신호로 취급하고, 그 크기 및 방향적 특성을 간결한 스펙트럼 통계로 요약하며, 과거 이미지-질문 쌍을 재생하거나 생성된 가짜 예제 또는 이전 단계의 모델 스냅샷을 저장하는 대신 기술별 프로토타입 분포만 저장합니다. 후속 단계에서 위상 불변 스펙트럼 정규화기는 이러한 프로토타입의 유해한 드리프트를 방지하면서 동시에 인스턴스 수준의 어텐션이 새로운 작업에 적응할 수 있도록 제약합니다. 우리는 기술 기반 스펙트럼 드리프트가 스펙트럼 충분성 가정 하에서 잊힘을 제어하고, 푸리에 전력 스펙트럼이 공간 이동 및 경계 내의 작은 변화에 안정적임을 보여주는 이론적 분석을 제공합니다. VQA v2, VQACL, CLT-VQA, CoIN 및 UCIT를 포함한 지속적인 VQA 및 멀티모달 명령어 튜닝 벤치마크에서 실험 결과, ASR은 강력한 재생 기반, 정규화 기반 및 어댑터 기반 모델보다 최종 성능을 일관되게 향상시키고 잊힘을 줄입니다. 기술 수준의 어텐션 구조를 보존하는 것은 지속적인 MLLM을 위한 효과적이고 가벼운 메커니즘입니다. 코드는 https://github.com/Creative-zcx/attention-spectrum-replay 에서 확인할 수 있습니다.

Original Abstract

Multimodal large language models (MLLMs) are increasingly required to adapt to non-stationary streams of visual domains, question types, and user instructions, yet continual fine-tuning often causes severe forgetting of previously acquired multimodal skills. Existing continual vision-language methods mainly preserve outputs, replay data or pseudo-data, regularize embedding geometry, or allocate task-specific parameters, but they provide limited control over how internal cross-modal attention patterns supporting old skills drift during adaptation. We propose Attention-Spectrum Regularization (ASR), a replay-free continual learning framework that preserves skill-conditioned structures of cross-modal attention. ASR treats cross-attention maps as two-dimensional signals, summarizes their scale and directional properties into compact spectral statistics, and stores only skill-wise prototype distributions instead of replaying past image-question pairs, generated pseudo-examples, or old-stage teacher snapshots. In later stages, a phase-invariant spectral regularizer constrains harmful drift of these prototypes while allowing instance-level attention to adapt to new tasks. We provide theoretical analysis showing that skill-conditioned spectral drift controls forgetting under a spectral sufficiency assumption, and that Fourier power spectra are stable to spatial translations and bounded perturbations. Experiments on continual VQA and multimodal instruction-tuning benchmarks, including VQA v2, VQACL, CLT-VQA, CoIN, and UCIT, show that ASR consistently improves final performance and reduces forgetting over strong replay-, regularization-, and adapter-based baselines. Preserving skill-level attention structure is an effective and lightweight mechanism for continual MLLMs. Code is available at https://github.com/Creative-zcx/attention-spectrum-replay

1 Citations
0 Influential
30.9657359028 Altmetric
6.9 Score
Original PDF
1

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!