모든 이미지는 위험한 이야기를 담고 있다: 메모리 기반 다중 에이전트 공격을 통한 VLM 제어 우회 연구
Every Picture Tells a Dangerous Story: Memory-Augmented Multi-Agent Jailbreak Attacks on VLMs
비전-언어 모델(VLM)의 빠른 발전은 인공지능 분야에 획기적인 능력을 제공했지만, 이러한 지속적인 모달 확장으로 인해 훨씬 더 넓고 제약 없는 적대적 공격 표면이 노출되었습니다. 현재의 다중 모드 제어 우회 전략은 주로 표면 수준의 픽셀 변경, 타이포그래피 공격 또는 악성 이미지에 초점을 맞추지만, 시각 데이터에 내재된 복잡한 의미 구조와 상호 작용하지 못합니다. 이는 원본의 자연스러운 이미지에서 잠재적인 의미 기반 공격 표면이 대부분 검토되지 않고 있음을 의미합니다. 이러한 깊은 의미적 취약점을 드러내기 위해, 우리는 시각적 의미를 명시적으로 활용하여 자동화된 제어 우회 공격을 수행하는 메모리 기반 다중 에이전트 제어 우회 프레임워크인 **MemJack**을 소개합니다. MemJack은 조정된 다중 에이전트 협력을 통해 시각적 개체를 악의적인 의도로 연결하고, 다각적인 시각-의미적 위장을 통해 적대적 프롬프트를 생성하며, 반복적 Nullspace Projection (INLP) 기하학적 필터를 사용하여 조기 잠재 공간 거부를 우회합니다. MemJack은 지속적인 다중 모드 경험 메모리를 통해 성공적인 전략을 축적하고 이전하여, 다양한 이미지에서 일관성 있는 다중 턴 제어 우회 공격 상호 작용을 유지함으로써 새로운 이미지에 대한 공격 성공률(ASR)을 향상시킵니다. 완전하고 수정되지 않은 COCO val2017 이미지에 대한 광범위한 실험 결과, MemJack은 Qwen3-VL-Plus에서 71.48%의 ASR을 달성했으며, 확장된 예산 하에서는 90%까지 성능이 향상됩니다. 또한, 향후 방어적 정렬 연구를 촉진하기 위해, 113,000개 이상의 상호 작용형 다중 모드 제어 우회 공격 경로를 포함하는 포괄적인 데이터셋인 **MemJack-Bench**를 공개하여, 본질적으로 강력한 VLM을 개발하기 위한 중요한 기반을 제공합니다.
The rapid evolution of Vision-Language Models (VLMs) has catalyzed unprecedented capabilities in artificial intelligence; however, this continuous modal expansion has inadvertently exposed a vastly broadened and unconstrained adversarial attack surface. Current multimodal jailbreak strategies primarily focus on surface-level pixel perturbations and typographic attacks or harmful images; however, they fail to engage with the complex semantic structures intrinsic to visual data. This leaves the vast semantic attack surface of original, natural images largely unscrutinized. Driven by the need to expose these deep-seated semantic vulnerabilities, we introduce \textbf{MemJack}, a \textbf{MEM}ory-augmented multi-agent \textbf{JA}ilbreak atta\textbf{CK} framework that explicitly leverages visual semantics to orchestrate automated jailbreak attacks. MemJack employs coordinated multi-agent cooperation to dynamically map visual entities to malicious intents, generate adversarial prompts via multi-angle visual-semantic camouflage, and utilize an Iterative Nullspace Projection (INLP) geometric filter to bypass premature latent space refusals. By accumulating and transferring successful strategies through a persistent Multimodal Experience Memory, MemJack maintains highly coherent extended multi-turn jailbreak attack interactions across different images, thereby improving the attack success rate (ASR) on new images. Extensive empirical evaluations across full, unmodified COCO val2017 images demonstrate that MemJack achieves a 71.48\% ASR against Qwen3-VL-Plus, scaling to 90\% under extended budgets. Furthermore, to catalyze future defensive alignment research, we will release \textbf{MemJack-Bench}, a comprehensive dataset comprising over 113,000 interactive multimodal jailbreak attack trajectories, establishing a vital foundation for developing inherently robust VLMs.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.