시각-언어 모델의 취약성 분석: 텍스처 제약 기반 교란 및 다중 모드 최적화를 통한 다중 모드 적대적 시너지
Unveiling the Fragility of Vision-Language Models: Multi-Modal Adversarial Synergy via Texture-Constrained Perturbations and Cross-Modal Optimization
대규모 시각-언어 모델(LVLM)은 이미지 설명 생성 및 시각 질의 응답과 같은 다양한 작업에서 시각 및 텍스트 입력을 통합하여 다중 모드 이해 분야에 혁신을 가져왔습니다. 그러나 특히 두 가지 모드를 모두 활용하는 적대적 공격에 대한 LVLM의 견고성은 아직 충분히 연구되지 않았으며, 이는 자율 주행 및 콘텐츠 검열과 같은 중요한 응용 분야에 위험을 초래할 수 있습니다. 기존 공격은 단일 모드에 집중하거나 비현실적인 화이트박스 접근 방식을 요구하여 실제 적용 가능성이 제한됩니다. 본 논문에서는 LVLM에 대한 보편적이고 블랙박스 다중 모드 공격을 설계하는 획기적인 프레임워크인 Multi-Modal Adversarial Synergy(MMAS)를 소개합니다. MMAS는 모델 쿼리만을 사용하여 이미지에 대한 텍스처 스케일 제약 조건 기반의 보편적 적대적 교란과 텍스트에 대한 학습 가능한 프롬프트 교란을 동시에 생성하며, 이 두 가지 교란은 공동으로 최적화됩니다. 이미지 교란은 웨이블릿 기반 텍스처 제약을 활용하여 다양한 시각 입력에 대한 인지 불가능성과 견고성을 보장합니다. 텍스트 교란은 임베딩 공간에서의 L-norm 제약을 통해 의미론적 일관성을 유지하면서 출력을 원하는 방향으로 유도합니다. 새로운 크로스 모드 정규화 항은 교란의 기울기 방향을 조정하여 시너지 효과를 높이고 작업 및 모델 간의 전이성을 향상시킵니다. 광범위한 실험 결과, 제안하는 공격 방법이 다양한 LVLM에서 강력한 보편적 적대적 능력을 갖는다는 것을 보여줍니다.
Large Vision-Language Models (LVLMs) have transformed multi-modal understanding, excelling in tasks like image captioning and visual question answering by integrating visual and textual inputs. However, their robustness against adversarial attacks, particularly those exploiting both modalities, remains underexplored, posing risks to critical applications like autonomous driving and content moderation. Existing attacks focus on single modalities or require impractical white-box access, limiting their real-world relevance. In this paper, we introduce Multi-Modal Adversarial Synergy, a groundbreaking framework that crafts universal, black-box multi-modal attacks against LVLMs. MMAS simultaneously generates a texture scale-constrained universal adversarial perturbation for images and a learnable prompt perturbation for text, optimized jointly using only model queries. The image perturbation leverages wavelet-based texture constraints to ensure imperceptibility and robustness across diverse visual inputs. The text perturbation, constrained by an L-norm in the embedding space, maintains semantic coherence while steering outputs toward a target. A novel cross-modal regularization term aligns the perturbations' gradient directions, enhancing their synergistic impact and transferability across tasks and models. Extensive experiments show the strong universal adversarial capabilities of our proposed attack with prevalent LVLMs.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.