2608.01849v1 Aug 03, 2026 cs.AI

미 학습된 멀티모달 대규모 언어 모델에서 지식 공백 탐색 및 연결

Exploring and Bridging Knowledge Holes in Unlearned Multimodal Large Language Models

Shu Wu
Shu Wu
Citations: 19
h-index: 2
Junxian You
Junxian You
Citations: 43
h-index: 3
Junkai Chen
Junkai Chen
Citations: 252
h-index: 8
Yuhao He
Yuhao He
Citations: 2
h-index: 1
Ruiqi Liu
Ruiqi Liu
Citations: 83
h-index: 2
Zhetao Guo
Zhetao Guo
Citations: 0
h-index: 0

머신 언러닝은 멀티모달 대규모 언어 모델(MLLM)에서 유해한 콘텐츠를 제거하는 데 유망한 접근 방식이지만, 언러닝의 정확성을 보장하는 것은 여전히 중요한 과제입니다. 그 한 가지 이유는 현재 MLLM 언러닝 평가 패러다임이 중요한 결점을 가지고 있기 때문인데, 이는 벤치마크를 통해 모델의 유용성을 평가하지만, 이러한 벤치마크의 표현은 삭제 대상(forget set)과 거리가 멀어 '지식 공백'이라는 심각한 문제, 즉 무해한 관련 입력에 대한 성능 저하를 제대로 반영하지 못합니다. 본 연구에서는 언러닝된 MLLM에서 지식 공백을 탐색하기 위해, 삭제 대상과 유사한 일반적인 패턴을 공유하는 무해한 입력에 대한 의도치 않은 성능 저하를 포착하는 벤치마크를 구축하고, 통제된 실험을 통해 이러한 현상이 일반적으로 사용되는 접근 방식의 체계적인 결과임을 확인했습니다. 또한, 이러한 격차를 해소하기 위해, 본 연구에서는 '앵커드 정규화 기반 선택적 보호(Selective Protection with Anchored Regularization)'라는 방법을 제안합니다. 이 방법은 앵커드 활성화 필터링을 통해 일반적인 패턴을 보호하고, 엔티티 추상화를 통한 향상을 통해 이를 강화합니다. SafeEraser에 대한 실험 결과, SPAR은 표준 baseline보다 훨씬 우수한 성능(98% 이상)으로 원래 응답 품질을 복구했으며, 동시에 공격 성공률 0%, 경쟁력 있는 모델 유용성을 달성했습니다. 이러한 결과는 신뢰할 수 있는 MLLM 언러닝을 위한 더욱 세밀한 평가의 필요성을 강조합니다.

Original Abstract

Machine unlearning offers a promising approach to remove unsafe content from Multimodal Large Language Models (MLLMs), yet ensuring the precision of unlearning remains a persistent challenge. One reason is that current MLLM unlearning evaluation paradigms suffer from a critical blind spot: they assess model utility through benchmarks whose representations are distant from the forget set, failing to capture knowledge holes---severe degradation on benign adjacent inputs. To probe knowledge holes in unlearned MLLMs, we construct a benchmark that captures unintended degradation on benign inputs sharing generic patterns with the forget set, and confirm through controlled experiments that they are a systematic consequence of commonly used approaches. Furthermore, to bridge this gap, we propose Selective Protection with Anchored Regularization, which protects generic patterns via anchored activation filtering while reinforcing them through entity-abstracted enhancement. Our experiments on SafeEraser demonstrate that SPAR recovers over 98% of vanilla response quality compared to below 50% for standard baselines---while achieving 0.00% attack success rate and competitive model utility. These results underscore the necessity of more fine-grained evaluation for trustworthy MLLM unlearning.

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!