2605.05709v1 May 07, 2026 cs.AI

숨기기, 재구성, 탈출: MLLM에서 재구성-숨기기 균형을 이용한 공격

Conceal, Reconstruct, Jailbreak: Exploiting the Reconstruction-Concealment Tradeoff in MLLMs

Richeng Jin
Richeng Jin
Citations: 746
h-index: 14
Huaiyu Dai
Huaiyu Dai
Citations: 16
h-index: 2
Md Farhamdur Reza
Md Farhamdur Reza
Citations: 98
h-index: 5
Tianfu Wu
Tianfu Wu
Citations: 8
h-index: 2

다중 모드 대규모 언어 모델(MLLM)에 대한 의도-가림 기술 기반 탈출 공격은 유해한 쿼리를 안전 장치를 우회할 수 있는 숨겨진 다중 모드 입력으로 변환합니다. 우리는 이러한 공격이 extit{재구성-숨기기 균형}에 의해 지배된다는 것을 보여줍니다. 즉, 변환된 입력은 안전 필터로부터 유해한 의도를 숨기면서 동시에 대상 모델이 원래 요청을 재구성할 수 있을 만큼 충분히 복구 가능해야 합니다. 세 가지 대표적인 블랙박스 방법에 대한 재구성 분석을 통해, 기존 변환 방법은 이 균형을 맞추는 데 어려움을 겪으며, 이는 그 효과를 제한합니다. 반대로, 문자 제거 변형은 더 나은 균형을 제공한다는 것을 보여줍니다. 이를 바탕으로, 우리는 extit{가림-인식 변형 생성} 방법을 제안합니다. 이 방법은 유해 키워드와의 일치도가 낮고 서로 다양한 문자 제거 변형을 선택하고, 다섯 가지 모드 인식 프롬프트 전략을 통해 이를 구현합니다. 또한, 우리는 유해 키워드를 다양한 맥락에서 묘사하는 extit{키워드 관련 방해 이미지}를 도입하여, 일반적인 방해 이미지보다 효과적인 보조 시각적 맥락을 제공합니다. 폐쇄형 및 오픈 소스 MLLM에 대한 실험 결과, 제안된 전략이 강력한 기준 방법을 능가하며, 모델 자체의 재구성 능력이 악용되어 숨겨진 유해한 의도를 복구하고 안전하지 않은 응답을 생성할 수 있다는 미개척된 취약점을 드러냅니다.

Original Abstract

Intent-obfuscation-based jailbreak attacks on multimodal large language models (MLLMs) transform a harmful query into a concealed multimodal input to bypass safety mechanisms. We show that such attacks are governed by a \emph{reconstruction--concealment tradeoff}: the transformed input must hide harmful intent from safety filters while remaining recoverable enough for the victim model to reconstruct the original request. Through a reconstruction analysis of three representative black-box methods, we find that existing transformations struggle to balance this tradeoff, limiting their effectiveness. In contrast, we show that character-removed variants achieve a better balance. Building on this, we propose \emph{concealment-aware variant construction}, which greedily selects character-removed variants that are low in harmful-keyword alignment and mutually diverse, and instantiates them through five modality-aware prompting strategies. We further introduce \emph{keyword-related distractor images} that depict the harmful keyword in diverse contexts, providing more effective auxiliary visual context than generic distractor images. Experiments across closed-source and open-source MLLMs show the proposed strategies outperform strong baselines, revealing an underexplored vulnerability: a model's own reconstruction ability can be exploited to recover hidden harmful intent and produce unsafe responses.

0 Citations
0 Influential
7 Altmetric
35.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!