2608.11691v1 Aug 12, 2026 cs.LG

LEMUR: 시각적 앵커 기반 추론 방향 전환을 통한 잠재 엔트로피 인지 다중 모드 모델의 학습 제거

LEMUR: Latent Entropy-aware Multimodal Unlearning via Visual-anchored Reasoning Redirection

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Hao Fang
Hao Fang
Citations: 315
h-index: 10
Xinhao Zhong
Xinhao Zhong
Citations: 63
h-index: 5
Yu Qiao
Yu Qiao
Citations: 3,128
h-index: 12
Junhao Li
Junhao Li
Citations: 19
h-index: 3

강화 학습(RL) 후속 훈련은 멀티모달 대규모 추론 모델(MLRM)에 탐색적인 사고 과정(CoT)을 부여하여 시각적 추론 능력을 크게 향상시킵니다. 그러나 저희는 이러한 기능이 다음과 같은 특정한 개인 정보 보호 취약점을 야기한다는 것을 발견했습니다. 즉, 민감한 사실이 최종 답변에서 성공적으로 제거되더라도 모델은 여전히 추론 과정에서 이를 재현할 수 있습니다. 이와 같은 정보 유출은 강화 학습을 통해 훈련된 원시 MLRM에서 기존의 비-추론 기반 모델보다 훨씬 더 두드러지며, 이는 기존의 학습 제거 방법으로는 해결할 수 없는 개인 정보 보호 위험을 드러냅니다. 저희는 강화 학습으로 인한 탐색 과정이 민감한 콘텐츠에 독특한 토큰 수준의 엔트로피 특징을 남기는데, 이 특징은 원시 모델에서는 거의 나타나지 않는다는 것을 확인했습니다. 이러한 관찰을 바탕으로, 저희는 원시적으로 강화 학습을 통해 훈련된 멀티모달 모델을 위한 완전한 학습 제거 방식인 LEMUR 프레임워크를 제안합니다. LEMUR은 엔트로피 변화를 제어 신호로 사용하여 민감한 추론이 시작되는 시점과 삭제 과정을 중단해야 하는 시점을 식별합니다. 이 과정에서 LEMUR은 엔트로피 변조된 시각적 앵커를 활용하여 추론 경로를 재지향하고, 입력 이미지에 다시 연결된 정제된 확률 가중 임베딩으로 대체된 토큰을 사용합니다. 다양한 MLRM 모델에서 LEMUR은 기존의 학습 제거 방법보다 추론 과정 및 답변 유출을 효과적으로 억제하는 동시에 비-민감한 유용성과 출력의 자연스러움을 더 잘 유지합니다. 이러한 결과는 강화 학습으로 인한 엔트로피 변화가 개인 정보 보호 유출에 대한 독특한 신호를 제공하며, 이러한 신호를 활용하면 추론 능력을 갖춘 멀티모달 모델을 위한 효과적인 학습 제거를 수행할 수 있음을 보여줍니다.

Original Abstract

Reinforcement-learning (RL) post-training equips multimodal large reasoning models (MLRMs) with exploratory chains of thought (CoT), substantially improving visual reasoning. However, we find that this capability introduces a distinct privacy vulnerability: even when a sensitive fact is successfully unlearned from the final answer, the model may still reproduce it in its reasoning trace. This leakage is substantially more pronounced in natively RL-trained MLRMs than in their non -reasoning base models, revealing a privacy risk that existing unlearning methods are not designed to address. We show that RL-induced exploration leaves sensitive content with a distinctive token-level entropy signature that is largely absent from base models. Based on this observation, we propose LEMUR, a fully training-free, inference-time unlearning framework for natively RL-trained multimodal models. LEMUR uses entropy dynamics as a control signal to identify when sensitive reasoning begins and when sanitization should stop. During this interval, it redirects the reasoning trajectory through entropy-modulated visual-anchor latent injection, replacing committed tokens with sanitized, probability-weighted embeddings re-grounded in the input image. Across diverse MLRMs, LEMUR consistently outperforms existing unlearning met hods in suppressing both reasoning-trace and answer leakage, while better preserving non-sensitive utility and output fluency. These results demonstrate that RL-induced entropy dynamics provide a distinctive signal for privacy leakage and that exploiting this signal enables effective training-free unlearning for reasoning-capable multimodal models.

0 Citations
0 Influential
6 Altmetric
30.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!