2603.21697v1 Mar 23, 2026 cs.CR

구조화된 시각적 내러티브가 다중 모드 대규모 언어 모델의 안전성 정렬을 저해함

Structured Visual Narratives Undermine Safety Alignment in Multimodal Large Language Models

Yujia Hu
Yujia Hu
Citations: 77
h-index: 3
Roy Ka-wei Lee
Roy Ka-wei Lee
Citations: 208
h-index: 4
Rui Yang Tan
Rui Yang Tan
Citations: 55
h-index: 3

다중 모드 대규모 언어 모델(MLLM)은 시각적 추론 기능을 추가하여 텍스트 기반 LLM을 확장하지만, 동시에 시각적으로 기반한 지시 하에 새로운 안전성 문제를 야기합니다. 본 연구에서는 유해한 목표를 단순한 세 패널 시각적 내러티브 내에 숨겨 모델이 역할을 수행하고 "만화 완성"하도록 유도하는 만화 템플릿 기반의 공격을 분석합니다. JailbreakBench 및 JailbreakV를 기반으로, 우리는 10가지 유해 범주와 5가지 작업 환경에 걸쳐 1,167개의 공격 사례를 포함하는 만화 기반 공격 벤치마크인 ComicJailbreak를 소개합니다. 15개의 최첨단 MLLM(상용 모델 6개, 오픈 소스 모델 9개)을 대상으로 한 실험 결과, 만화 기반 공격은 강력한 규칙 기반 공격과 유사한 성공률을 보이며, 일반 텍스트 및 무작위 이미지 기반 공격보다 훨씬 뛰어난 성능을 보였습니다. 일부 상용 모델의 경우, 앙상블 공격 성공률이 90%를 초과했습니다. 또한, 기존의 방어 방법이 유해한 만화에 효과적이지만, 무해한 프롬프트에 사용될 경우 높은 거부율을 유발한다는 것을 확인했습니다. 마지막으로, 자동 평가 및 표적 인간 평가를 통해 현재의 안전성 평가 도구가 민감하지만 유해하지 않은 콘텐츠에 대해 신뢰성이 낮을 수 있음을 보여줍니다. 본 연구 결과는 내러티브 기반의 다중 모드 공격에 대한 안전성 정렬의 중요성을 강조합니다.

Original Abstract

Multimodal Large Language Models (MLLMs) extend text-only LLMs with visual reasoning, but also introduce new safety failure modes under visually grounded instructions. We study comic-template jailbreaks that embed harmful goals inside simple three-panel visual narratives and prompt the model to role-play and "complete the comic." Building on JailbreakBench and JailbreakV, we introduce ComicJailbreak, a comic-based jailbreak benchmark with 1,167 attack instances spanning 10 harm categories and 5 task setups. Across 15 state-of-the-art MLLMs (six commercial and nine open-source), comic-based attacks achieve success rates comparable to strong rule-based jailbreaks and substantially outperform plain-text and random-image baselines, with ensemble success rates exceeding 90% on several commercial models. Then, with the existing defense methodologies, we show that these methods are effective against the harmful comics, they will induce a high refusal rate when prompted with benign prompts. Finally, using automatic judging and targeted human evaluation, we show that current safety evaluators can be unreliable on sensitive but non-harmful content. Our findings highlight the need for safety alignment robust to narrative-driven multimodal jailbreaks.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!