VisInject: 방해 ≠ 주입 -- 비전-언어 모델에 대한 범용 적대적 공격의 이중 차원 평가
VisInject: Disruption != Injection -- A Dual-Dimension Evaluation of Universal Adversarial Attacks on Vision-Language Models
정렬된 다중 모달 대규모 언어 모델에 대한 범용 적대적 공격이 점점 더 많이 보고되고 있으며, 공격 성공률이 60-80%에 달하는 것으로 나타납니다. 이는 시각 모달리티가 프롬프트 주입 채널로서 미세한 변화에 매우 취약하다는 것을 시사합니다. 본 연구에서는 이 수치가 두 가지 상이한 현상을 혼동한다고 주장합니다. 즉, (i) 모델의 출력이 변경되었는지 (영향), 그리고 (ii) 공격자가 선택한 특정 개념이 실제로 출력되었는지 (정밀 주입)입니다. 우리는 기존의 두 가지 기술인 Universal Adversarial Attack과 AnyAttack을 $L_{inf}$ 예산 16/255 하에서 결합하고, 이중 축 평가를 추가했습니다. 첫 번째 축은 프로그램 기반의 Ratcliff-Obershelp 드리프트 점수를 사용하여 '영향'을 평가하고, 두 번째 축은 4단계의 순위 범주(none/weak/partial/confirmed)를 사용하여 '정밀 주입'을 평가합니다. 판별기로 DeepSeek-V4-Pro를 사용하며, thinking 모드로 작동하고 Claude Opus 4.7을 기준으로 교정되었습니다. Cohen's $κ$ 값은 주입 축에서 0.77으로, 상당히 높은 수준의 일치도를 나타냅니다. 전체 4475개의 입력 데이터는 SHA-256 해시로 저장되어 있으며, 데이터셋과 함께 제공되므로 검토자가 API 키 없이도 논문의 결과를 재현할 수 있습니다. 네 가지 개방형 비전-언어 모델, 일곱 가지 공격 프롬프트, 그리고 일곱 가지 테스트 이미지를 사용하여 총 6615개의 쌍을 분석한 결과, 두 축은 약 90배의 차이를 보였습니다. 66.4%의 쌍에서 프로그램적으로 출력이 변경되었지만 (LLM 판별 결과, 상당하거나 완전한 수준으로 평가된 46.6%), 실제로 '정밀 주입'이 발생한 경우는 0.756% (50/6615)에 불과하며, '정확히 일치하는' 경우는 0.030% (2/6615)에 불과했습니다. 실제로 주입이 발생한 경우는 주로 스크린샷 또는 문서 형태의 이미지에서 나타났으며, 이러한 이미지의 의미 자체가 텍스트 변환을 유도하는 경향이 있습니다. BLIP-2 모델은 모든 2205개의 쌍에서 $L_{inf}$ = 16/255일 때 '영향'에 대한 검출 가능한 드리프트가 extit{하나도 발생하지 않았습니다}, 심지어 Stage-1의 대체 모델로 사용했을 때조차도 그렇습니다. 본 연구에서는 전체 데이터셋(21개의 범용 이미지, 147개의 적대적 이미지, 6615개의 응답 쌍, v3의 이중 축 판별 결과, 그리고 캐시)을 huggingface.co/datasets/jeffliulab/visinject에서 공개합니다.
Universal adversarial attacks on aligned multimodal large language models are increasingly reported with attack success rates in the 60-80% range, suggesting the visual modality is highly vulnerable to imperceptible perturbations as a prompt-injection channel. We argue that this number conflates two distinct events: (i) the model's output was perturbed (Influence), and (ii) the attacker's chosen target concept was actually emitted (Precise Injection). We compose two existing techniques -- Universal Adversarial Attack and AnyAttack -- under an $L_{inf}$ budget of 16/255, and we add a dual-axis evaluation: a deterministic Ratcliff-Obershelp drift score for Influence (programmatic baseline) plus a 4-tier ordinal categorical none/weak/partial/confirmed for Precise Injection. The judge is DeepSeek-V4-Pro in thinking mode, calibrated against Claude Opus 4.7 with Cohen's $κ$ = 0.77 on the injection axis (substantial agreement); the entire 4475-entry SHA-256 input cache ships with the dataset so reviewers can re-derive paper numbers bit-exact without an API key. Across 6615 pairs over four open VLMs, seven attack prompts, and seven test images, the two axes diverge by roughly 90$\times$: 66.4% of pairs are programmatically disturbed (LLM-judged 46.6% at the substantial-or-complete tier), but only 0.756% (50/6615) reach any non-none injection tier and only 0.030% (2/6615) verbatim. The few injections that do land cluster on screenshot- or document-style carriers whose semantics already invite text transcription. BLIP-2 shows \emph{zero detectable drift} at $L_{inf}$ = 16/255 across all 2205 pairs even when used as a Stage-1 surrogate. We release the full dataset -- 21 universal images, 147 adversarial photos, 6,615 response pairs, the v3 dual-axis judge results, and the cache at huggingface.co/datasets/jeffliulab/visinject.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.