시각-언어 모델의 적대적 강건성 전문가 통합
Unifying Adversarially Robust Model Experts in Vision-Language Models
CLIP과 같은 시각-언어 모델(VLM)은 실제 응용 및 배포에 심각한 문제를 야기하는 적대적 공격에 취약합니다. 적대적 미세 조정은 주요 방어 방법으로 부상하고 있지만, 다양한 미세 조정 전략은 종종 뚜렷한 강건성 특성을 가진 전문화된 모델을 생성합니다. 각 미세 조정된 모델은 특정 평가 환경에서는 뛰어난 성능을 보이지만 다른 환경에서는 그렇지 못하여 전반적인 방어 능력을 제한합니다. 이러한 전문화된 미세 조정된 모델을 '강건성 전문가'라고 부르며, 우리는 임베딩 공간 정렬을 사용하는 협력적 적대적 강건성 미세 조정 프레임워크인 CARE(Collaborative Adversarial Robustness fine-tuning)를 제안합니다. CARE는 훈련 중에 여러 전문가를 유지하고, 임베딩 공간의 조화를 통해 지식 교환을 가능하게 하며, 학습된 지식을 단일 통합된 강건성 모델로 통합합니다. 각 전문가는 서로에게 이익을 얻으면서도 개별적인 특성을 유지하여 최종 모델이 상호 보완적인 강건성 속성을 갖도록 합니다. 본 논문에서는 상호 보완적인 강건성 동작을 가진 두 가지 다른 적대적 미세 조정 전략에 CARE를 적용하여 그 효과를 입증합니다. 기존 이미지 분류 및 다양한 시각-언어 작업에 대한 광범위한 실험 결과, 우리의 접근 방식인 CARE가 개별적으로 학습된 모델 전문가보다 우수한 성능을 보이는 것을 확인했습니다. 이러한 결과는 모델 전문가 간의 협력적 학습이 적대적 강건성을 향상시키는 유망한 방향임을 시사합니다.
Vision-language models (VLMs), such as CLIP, are vulnerable to adversarial attacks, posing a serious problem for real-life applications and deployment. Adversarial fine-tuning emerges as a prominent defense method; however, different fine-tuning strategies often produce specialized models with distinct robustness characteristics. Each fine-tuned model in turn thrives in some evaluation settings but falters on others, limiting their defensive capabilities. We refer to these specialized fine-tuned models as robust model experts and propose a collaborative adversarial fine-tuning framework: CARE - Collaborative Adversarial Robustness fine-tuning using Embedding alignment. CARE maintains multiple experts during training, enables knowledge exchange through embedding-space harmonization, and consolidates the learned knowledge into a single unified robust model. Experts benefit from one another while preserving their individual specializations, enabling the final model to inherit complementary robustness properties. In this paper, we demonstrate CARE on two different adversarial fine-tuning strategies with complementary robustness behaviors. Extensive experiments on classic image classification and downstream vision-language tasks display the effectiveness of our approach, with CARE being able to outperform individually learned model experts. The results suggest that collaborative learning across model experts is a promising direction for improving adversarial robustness.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.