지속적인 다중 모드 학습에서의 숨겨진 망각: 정확도는 유지되지만 연관성이 손실되는 경우
Hidden Forgetting in Continual Multimodal Learning: When Accuracy Survives but Grounding Fails
다중 모드 대규모 언어 모델은 끊임없이 변화하는 작업과 영역에 적응해야 하지만, 기존의 지속적 학습 지표는 주로 과거 답변의 정확성을 측정하며, 다중 모드 연관성의 안정성은 대부분 간과됩니다. 본 연구에서는 이러한 간과된 실패 양상을 분석하고, 지속적으로 적응된 다중 모드 언어 모델이 답변 내용뿐만 아니라 시각, 텍스트, OCR, 차트 및 문서 증거를 사용하는 방식까지 유지할 수 있는지 질문합니다. 우리는 '숨겨진 증거 사용 망각(hidden evidence-use forgetting)' 현상을 밝혀내는데, 이는 답변 정확도는 유지되지만 모델이 조용히 다른 또는 연관성이 낮은 증거 채널로 이동하는 현상입니다. 이를 해결하기 위해, 리플레이 방식 없이 의존성 제약을 적용한 지속적 학습 프레임워크인 extsc{RCL}을 제안합니다. extsc{RCL}은 이전 체크포인트를 행동 기준점으로 고정하고, 반사실적 채널 개입을 통해 교사 및 학생 모델의 증거 의존성 프로필을 추정하며, 작업 학습, 예측 보존, 그리고 의존성 보존을 동시에 최적화합니다. CoIN, COAST, MCITlib 데이터셋과 연관성 민감한 다중 모드 스트림에서 extsc{RCL}은 리플레이 방식, PEFT(Parameter-Efficient Fine-Tuning), 라우팅, 메모리 지원 방식을 사용하는 기존 방법보다 최종 성능을 향상시키고 망각 현상을 줄이며, 동시에 모달리티 의존성 변화, 주요 증거 채널 변경 및 숨겨진 망각 비율을 크게 감소시킵니다. 이러한 결과는 견고한 지속적인 다중 모드 학습을 위해서는 정확한 답변뿐만 아니라, 그 답변에 대한 증거 경로까지 유지하는 것이 중요하다는 것을 시사합니다.
Multimodal large language models must continually adapt to evolving tasks and domains, yet standard continual learning metrics mainly measure whether old answers remain correct, leaving the stability of multimodal grounding largely unexamined. We study this overlooked failure mode and ask whether a continually adapted MLLM can preserve not only what it answers, but also how it uses visual, textual, OCR, chart, and document evidence. We identify \emph{hidden evidence-use forgetting}, where answer accuracy is retained while the model silently shifts toward different or less grounded evidence channels, and propose \textsc{RCL}, a replay-free reliance-constrained continual learning framework. \textsc{RCL} freezes the previous checkpoint as a behavioral reference, estimates teacher and student evidence-reliance profiles through counterfactual channel interventions, and jointly optimizes task learning, prediction preservation, and reliance preservation without adding inference-time cost. Across CoIN, COAST, MCITlib, and an evidence-sensitive multimodal stream, \textsc{RCL} consistently improves final performance and reduces forgetting over replay-free, PEFT, routing, and memory-assisted baselines, while substantially lowering modality reliance drift, dominant evidence flips, and hidden forgetting rates. These results suggest that robust continual multimodal learning requires preserving the evidence path behind correct answers, not merely the answers themselves.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.