2603.29418v1 Mar 31, 2026 cs.CV

다중 모드 대규모 언어 모델에 대한 적대적 프롬프트 주입 공격

Covert Visual Prompt Injection against Commercial Multimodal Large Language Models

Meiwen Ding
Meiwen Ding
Citations: 26
h-index: 2
Song Xia
Song Xia
Citations: 62
h-index: 4
Chenqi Kong
Chenqi Kong
Citations: 240
h-index: 7
Xudong Jiang
Xudong Jiang
Citations: 44
h-index: 2

다중 모드 대규모 언어 모델(MLLM)이 실제 응용 분야에 점점 더 많이 사용되고 있지만, 지시사항을 따르는 특성으로 인해 프롬프트 주입 공격에 취약합니다. 기존의 프롬프트 주입 방법은 주로 텍스트 프롬프트 또는 인간 사용자가 인지할 수 있는 시각적 프롬프트에 의존합니다. 본 연구에서는 강력한 비공개 MLLM에 대한 인지 불가능한 시각적 프롬프트 주입 공격을 연구합니다. 여기에서 적대적인 지시사항은 시각적 모드에 내장됩니다. 저희 방법은 악성 프롬프트를 입력 이미지에 제한된 텍스트 오버레이를 통해 적응적으로 내장하여 의미론적 지침을 제공합니다. 동시에, 미세한 시각적 왜곡은 반복적으로 최적화되어 공격받는 이미지의 특징 표현을 악성 시각적 및 텍스트 목표와 거칠고 세밀한 수준 모두에서 일치시키도록 합니다. 특히, 시각적 목표는 텍스트 렌더링 이미지를 사용하여 구현되며, 최적화 과정에서 점진적으로 개선되어 원하는 의미를 더욱 정확하게 반영하고 전이성을 향상시킵니다. 두 가지 다중 모드 이해 작업에 대한 광범위한 실험 결과, 저희 방법이 기존 방법보다 우수한 성능을 보이는 것으로 나타났습니다.

Original Abstract

Although multimodal large language models (MLLMs) are increasingly deployed in real-world applications, their instruction-following behavior leaves them vulnerable to prompt injection attacks. Existing prompt injection methods predominantly rely on textual prompts or perceptible visual prompts that are observable by human users. In this work, we study imperceptible visual prompt injection against powerful closed-source MLLMs, where adversarial instructions are embedded in the visual modality. Our method adaptively embeds the malicious prompt into the input image via a bounded text overlay to provide semantic guidance. Meanwhile, the imperceptible visual perturbation is iteratively optimized to align the feature representation of the attacked image with those of the malicious visual and textual targets at both coarse- and fine-grained levels. Specifically, the visual target is instantiated as a text-rendered image and progressively refined during optimization to more faithfully represent the desired semantics and improve transferability. Extensive experiments on two multimodal understanding tasks across multiple closed-source MLLMs demonstrate the superior performance of our approach compared to existing methods.

1 Citations
1 Influential
3.5 Altmetric
20.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!