모든 것을 목표에 맞추세요: 멀티-크롭 라우팅 메타 최적화를 통한 폐쇄형 멀티모달 대규모 언어 모델 공격을 위한 범용 적대적 퍼터베이션
Make Anything Match Your Target: Universal Adversarial Perturbations against Closed-Source MLLMs via Multi-Crop Routed Meta Optimization
폐쇄형 멀티모달 대규모 언어 모델(MLLM)에 대한 표적 적대 공격은 블랙박스 전이 환경에서 점차 연구되고 있지만, 기존 방법은 주로 샘플별이며 입력 데이터 전반에 걸쳐 제한적인 재사용성을 제공합니다. 본 연구에서는 더욱 엄격한 환경인 범용 표적 전이 가능한 적대 공격(UTTAA)을 연구합니다. UTTAA는 단일 퍼터베이션을 사용하여 알려지지 않은 상업용 MLLM에서 임의의 입력을 특정 목표로 일관되게 유도해야 합니다. 기존의 샘플별 공격을 이러한 범용 환경으로 직접 적용하는 것은 세 가지 핵심적인 어려움에 직면합니다. (i) 표적 크롭의 무작위성으로 인해 목표 감독의 분산이 높아집니다. (ii) 토큰 단위의 매칭이 신뢰할 수 없는데, 이는 범용성이 이미지별 힌트를 억제하여 정렬을 고정시키는 역할을 하기 때문입니다. (iii) 적은 수의 소스 데이터를 사용하여 표적별로 최적화하는 과정은 초기화에 매우 민감하며, 이는 달성 가능한 성능을 저하시킬 수 있습니다. 본 연구에서는 멀티-크롭 집계와 어텐션 가이드 크롭을 통해 감독을 안정화하고, 정렬 가능성 게이팅된 토큰 라우팅을 통해 토큰 수준의 신뢰성을 향상시키며, 표적 간 퍼터베이션 사전 지식을 메타 학습하여 각 표적에 대한 더 강력한 솔루션을 제공하는 MCRMO-Attack을 제안합니다. 상업용 MLLM에서, 제안하는 방법은 GPT-4o에서 +23.7%, Gemini-2.0에서 +19.9%로, 가장 강력한 범용 기준 모델보다 공격 성공률을 향상시켰습니다.
Targeted adversarial attacks on closed-source multimodal large language models (MLLMs) have been increasingly explored under black-box transfer, yet prior methods are predominantly sample-specific and offer limited reusability across inputs. We instead study a more stringent setting, Universal Targeted Transferable Adversarial Attacks (UTTAA), where a single perturbation must consistently steer arbitrary inputs toward a specified target across unknown commercial MLLMs. Naively adapting existing sample-wise attacks to this universal setting faces three core difficulties: (i) target supervision becomes high-variance due to target-crop randomness, (ii) token-wise matching is unreliable because universality suppresses image-specific cues that would otherwise anchor alignment, and (iii) few-source per-target adaptation is highly initialization-sensitive, which can degrade the attainable performance. In this work, we propose MCRMO-Attack, which stabilizes supervision via Multi-Crop Aggregation with an Attention-Guided Crop, improves token-level reliability through alignability-gated Token Routing, and meta-learns a cross-target perturbation prior that yields stronger per-target solutions. Across commercial MLLMs, we boost unseen-image attack success rate by +23.7\% on GPT-4o and +19.9\% on Gemini-2.0 over the strongest universal baseline.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.