프롬프트가 시각적으로 변할 때: 대규모 이미지 편집 모델을 위한 시각 중심의 탈옥 공격
When the Prompt Becomes Visual: Vision-Centric Jailbreak Attacks for Large Image Editing Models
최근 대규모 이미지 편집 모델의 발전은 텍스트 기반 지침에서 시각 프롬프트 편집 패러다임으로 전환되었으며, 여기서 사용자 의도는 표식, 화살표 및 시각-텍스트 프롬프트와 같은 시각적 입력으로부터 직접 추론됩니다. 이러한 패러다임은 사용성을 크게 확장하지만, 중요한 안전 위험을 야기합니다. 즉, 공격 대상 자체가 시각적으로 변한다는 것입니다. 본 연구에서는 시각적 입력을 통해서만 악의적인 지시를 전달하는 최초의 시각-시각 탈옥 공격인 Vision-Centric Jailbreak Attack (VJA)을 제안합니다. 이 새로운 위협을 체계적으로 연구하기 위해, 이미지 편집 모델을 위한 안전 중심 벤치마크인 IESBench를 소개합니다. IESBench에 대한 광범위한 실험 결과, VJA는 최첨단 상용 모델을 효과적으로 공격하여 Nano Banana Pro에서 최대 80.9%, GPT-Image-1.5에서 70.1%의 공격 성공률을 달성했습니다. 이러한 취약점을 완화하기 위해, 우리는 추가적인 가드 모델 없이, 그리고 미미한 계산 오버헤드만 발생시키는, 자기 성찰적 다중 모드 추론 기반의 훈련이 필요 없는 방어 방법을 제안합니다. 이러한 방어 방법은 정렬되지 않은 모델의 안전성을 상용 시스템과 유사한 수준으로 향상시킵니다. 본 연구 결과는 새로운 취약점을 드러내며, 안전하고 신뢰할 수 있는 현대 이미지 편집 시스템을 발전시키기 위한 벤치마크와 실용적인 방어 방법을 제공합니다. 주의: 본 논문에는 대규모 이미지 편집 모델에 의해 생성된 불쾌한 이미지가 포함되어 있습니다.
Recent advances in large image editing models have shifted the paradigm from text-driven instructions to vision-prompt editing, where user intent is inferred directly from visual inputs such as marks, arrows, and visual-text prompts. While this paradigm greatly expands usability, it also introduces a critical and underexplored safety risk: the attack surface itself becomes visual. In this work, we propose Vision-Centric Jailbreak Attack (VJA), the first visual-to-visual jailbreak attack that conveys malicious instructions purely through visual inputs. To systematically study this emerging threat, we introduce IESBench, a safety-oriented benchmark for image editing models. Extensive experiments on IESBench demonstrate that VJA effectively compromises state-of-the-art commercial models, achieving attack success rates of up to 80.9% on Nano Banana Pro and 70.1% on GPT-Image-1.5. To mitigate this vulnerability, we propose a training-free defense based on introspective multimodal reasoning, which substantially improves the safety of poorly aligned models to a level comparable with commercial systems, without auxiliary guard models and with negligible computational overhead. Our findings expose new vulnerabilities, provide both a benchmark and practical defense to advance safe and trustworthy modern image editing systems. Warning: This paper contains offensive images created by large image editing models.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.