동적 방어 프로파일링을 통한 텍스트-이미지 모델의 인지적 탈옥
Dynamic Defense Profiling Enables Cognitive Jailbreak of Text-to-Image Models
텍스트-이미지(T2I) 생성 모델은 고품질 시각 콘텐츠 합성 분야에서 놀라운 발전을 이루었지만, 여전히 악의적인 사용에 취약하며, 특히 부적절한 이미지(NSFW)를 생성하는 데 이용될 수 있습니다. 대부분의 기존 탈옥 공격은 휴리스틱 기반 프롬프트 엔지니어링 또는 블랙박스 최적화에 의존하며, 모델 피드백을 성공 또는 실패라는 이분법적인 신호로 처리합니다. 이러한 단순화된 접근 방식은 텍스트 거부, 시각적 차단, 의미론적 필터링 등 다양한 실패 모드에 내재된 풍부한 정보를 간과하여 비효율적인 탐색과 심각한 의미론적 붕괴를 초래합니다. 본 논문에서는 MIND라는 인지적 탈옥 프레임워크를 제안합니다. MIND는 적대적 프롬프트 생성을 잠재적인 방어 메커니즘에 대한 믿음 상태 추론 문제로 재구성합니다. 기존 방법과 달리, MIND는 맹목적으로 우회 프롬프트를 찾는 대신, 다중 모드 피드백을 고밀도 신호로 해석하여 대상 시스템의 잠재적 방어 메커니즘을 적극적으로 모델링합니다. 구체적으로, 이 프레임워크는 세 가지 핵심 구성 요소로 이루어져 있습니다. (1) 미세한 피드백 분해를 위한 다중 모드 판별기(Multi-modal Judge), (2) 반복적인 믿음 상태 업데이트를 위한 방어 프로파일러(Defense Profiler), 그리고 (3) 과거에 효과적이었던 공격 전략을 검색하기 위한 메타 메모리(Meta-Memory) 모듈입니다. 이러한 구성 요소는 추론 기반의 진화 최적화 프로세스 내에서 통합되어, 적응적이고 의미적으로 일관된 탈옥 프롬프트 생성을 가능하게 합니다. I2P 벤치마크에 대한 광범위한 실험 결과, MIND는 효과적인 성능을 입증했습니다. Stable Diffusion v1.5 모델에 적용된 여섯 가지 대표적인 전처리 및 후처리 방어 설정 하에서, MIND는 95.62%의 공격 성공률(ASR)을 달성하여 기존 방법보다 훨씬 뛰어난 성능을 보였습니다. 또한, 제안된 프레임워크의 효과는 널리 사용되는 상용 T2I 시스템 네 가지에 대해 검증되었으며, Wan-2.5 모델에서 91.58%라는 최고 수준의 ASR을 달성했습니다.
Text-to-Image (T2I) generative models have achieved remarkable progress in synthesizing high-quality visual content, yet they remain vulnerable to adversarial misuse, particularly in generating Not-Safe-For-Work (NSFW) images. Most existing jailbreak attacks primarily rely on heuristic prompt engineering or black-box optimization, treating model feedback as a binary signal (success or failure). This coarse-grained paradigm overlooks the rich information embedded in diverse failure modes, such as textual refusal, visual blocking, and semantic sanitization, resulting in inefficient exploration and severe semantic collapse. In this paper, we propose MIND, a cognitive jailbreak framework that reframes adversarial prompt generation as a belief-state inference problem over latent defense mechanisms. Instead of blindly searching for bypass prompts, MIND actively models the target system's latent defense mechanisms by interpreting multi-modal feedback as high-density signals. Specifically, the framework integrates three core components: (1) a Multi-modal Judge for fine-grained feedback decomposition, (2) a Defense Profiler for iterative belief updating, and (3) a Meta-Memory module for retrieving historically effective attack strategies. These components are unified within a reasoning-driven evolutionary optimization process, enabling adaptive and semantically consistent jailbreak generation. Extensive experiments on the I2P benchmark demonstrate the effectiveness of MIND. Under six representative pre-processing and post-processing defense settings applied to the Stable Diffusion v1.5 model, MIND achieves an Attack Success Rate (ASR) of 95.62%, significantly outperforming existing methods. Additionally, the effectiveness of the proposed framework is validated across four widely used commercial T2I systems, achieving the highest ASR of 91.58% on Wan-2.5.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.