2605.01449v1 May 02, 2026 cs.CR

VisInject: 방해 ≠ 주입 -- 비전-언어 모델에 대한 범용 적대적 공격의 이중 차원 평가

VisInject: Disruption != Injection -- A Dual-Dimension Evaluation of Universal Adversarial Attacks on Vision-Language Models

Pangpang Liu
Pangpang Liu
Citations: 32
h-index: 3
Yingjie Lao
Yingjie Lao
Citations: 21
h-index: 2

정렬된 다중 모달 대규모 언어 모델에 대한 범용 적대적 공격이 점점 더 많이 보고되고 있으며, 공격 성공률이 60-80%에 달하는 것으로 나타납니다. 이는 시각 모달리티가 프롬프트 주입 채널로서 미세한 변화에 매우 취약하다는 것을 시사합니다. 본 연구에서는 이 수치가 두 가지 상이한 현상을 혼동한다고 주장합니다. 즉, (i) 모델의 출력이 변경되었는지 (영향), 그리고 (ii) 공격자가 선택한 특정 개념이 실제로 출력되었는지 (정밀 주입)입니다. 우리는 기존의 두 가지 기술인 Universal Adversarial Attack과 AnyAttack을 $L_{inf}$ 예산 16/255 하에서 결합하고, 이중 축 평가를 추가했습니다. 첫 번째 축은 프로그램 기반의 Ratcliff-Obershelp 드리프트 점수를 사용하여 '영향'을 평가하고, 두 번째 축은 4단계의 순위 범주(none/weak/partial/confirmed)를 사용하여 '정밀 주입'을 평가합니다. 판별기로 DeepSeek-V4-Pro를 사용하며, thinking 모드로 작동하고 Claude Opus 4.7을 기준으로 교정되었습니다. Cohen's $κ$ 값은 주입 축에서 0.77으로, 상당히 높은 수준의 일치도를 나타냅니다. 전체 4475개의 입력 데이터는 SHA-256 해시로 저장되어 있으며, 데이터셋과 함께 제공되므로 검토자가 API 키 없이도 논문의 결과를 재현할 수 있습니다. 네 가지 개방형 비전-언어 모델, 일곱 가지 공격 프롬프트, 그리고 일곱 가지 테스트 이미지를 사용하여 총 6615개의 쌍을 분석한 결과, 두 축은 약 90배의 차이를 보였습니다. 66.4%의 쌍에서 프로그램적으로 출력이 변경되었지만 (LLM 판별 결과, 상당하거나 완전한 수준으로 평가된 46.6%), 실제로 '정밀 주입'이 발생한 경우는 0.756% (50/6615)에 불과하며, '정확히 일치하는' 경우는 0.030% (2/6615)에 불과했습니다. 실제로 주입이 발생한 경우는 주로 스크린샷 또는 문서 형태의 이미지에서 나타났으며, 이러한 이미지의 의미 자체가 텍스트 변환을 유도하는 경향이 있습니다. BLIP-2 모델은 모든 2205개의 쌍에서 $L_{inf}$ = 16/255일 때 '영향'에 대한 검출 가능한 드리프트가 extit{하나도 발생하지 않았습니다}, 심지어 Stage-1의 대체 모델로 사용했을 때조차도 그렇습니다. 본 연구에서는 전체 데이터셋(21개의 범용 이미지, 147개의 적대적 이미지, 6615개의 응답 쌍, v3의 이중 축 판별 결과, 그리고 캐시)을 huggingface.co/datasets/jeffliulab/visinject에서 공개합니다.

Original Abstract

Universal adversarial attacks on aligned multimodal large language models are increasingly reported with attack success rates in the 60-80% range, suggesting the visual modality is highly vulnerable to imperceptible perturbations as a prompt-injection channel. We argue that this number conflates two distinct events: (i) the model's output was perturbed (Influence), and (ii) the attacker's chosen target concept was actually emitted (Precise Injection). We compose two existing techniques -- Universal Adversarial Attack and AnyAttack -- under an $L_{inf}$ budget of 16/255, and we add a dual-axis evaluation: a deterministic Ratcliff-Obershelp drift score for Influence (programmatic baseline) plus a 4-tier ordinal categorical none/weak/partial/confirmed for Precise Injection. The judge is DeepSeek-V4-Pro in thinking mode, calibrated against Claude Opus 4.7 with Cohen's $κ$ = 0.77 on the injection axis (substantial agreement); the entire 4475-entry SHA-256 input cache ships with the dataset so reviewers can re-derive paper numbers bit-exact without an API key. Across 6615 pairs over four open VLMs, seven attack prompts, and seven test images, the two axes diverge by roughly 90$\times$: 66.4% of pairs are programmatically disturbed (LLM-judged 46.6% at the substantial-or-complete tier), but only 0.756% (50/6615) reach any non-none injection tier and only 0.030% (2/6615) verbatim. The few injections that do land cluster on screenshot- or document-style carriers whose semantics already invite text transcription. BLIP-2 shows \emph{zero detectable drift} at $L_{inf}$ = 16/255 across all 2205 pairs even when used as a Stage-1 surrogate. We release the full dataset -- 21 universal images, 147 adversarial photos, 6,615 response pairs, the v3 dual-axis judge results, and the cache at huggingface.co/datasets/jeffliulab/visinject.

0 Citations
0 Influential
1.5 Altmetric
7.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!