AirflowAttack: 적외선 원격 감지 시각-언어 모델에 대한 열-공기 흐름 기반 적대적 공격
AirflowAttack: Thermal-Airflow Adversarial Perturbations against Infrared Remote-Sensing Vision-Language Models
시각-언어 모델(VLM)은 보안이 중요한 환경에서 점점 더 많이 적외선(IR) 원격 감지 이미지에 활용되고 있지만, 이들의 적대적 강건성은 아직 검토되지 않았습니다. 본 연구에서는 IR 원격 감지 VLM을 위한 최초의 공격 방법인 AirflowAttack을 제시하며, 이는 열-공기 흐름 난류를 적대적 교란의 기준으로 사용하는 첫 번째 시도입니다. 경량화된 생성기는 입력에 독립적인 단일 교란 패턴을 합성하며, 물리적으로 타당한 공기 흐름 패턴으로 정규화합니다. 하나의 대리 CLIP 모델로 최적화된 이 방법은 다섯 가지 다양한 CLIP 기반 모델에서 평균 48.5%의 제로샷 장면 분류 공격 성공률(ASR, 상위 1개 클래스가 변경되는 샘플 비율)을 달성하며, 이는 네 가지 IR 특화 물리적 기준(27.7--37.0%)보다 훨씬 높습니다. 이 방법은 최첨단 VLM 여섯 개에 적용되었으며, 장면 분류 정확도를 최대 38.2%까지 감소시켰지만, 역설적으로 일부 모델의 IR 분석 신뢰도를 높여 교란을 실제 열 증거(예: 온도 구배 및 대류)로 잘못 인식하게 만들었습니다. 추가 실험 결과, 공기 흐름 기준은 공격 성공률에 거의 영향을 미치지 않으면서 물리적 타당성을 향상시킵니다. 본 연구는 11개의 모델과 4가지 작업으로 구성된 벤치마크와 함께, 급속하게 확장되고 있는 IR VLM 생태계의 중요한 취약점을 드러냅니다.
Vision-language models (VLMs) are increasingly deployed on infrared (IR) remote sensing imagery in security-critical settings, yet their adversarial robustness remains unexamined. We present AirflowAttack, to our knowledge the first adversarial attack for IR remote-sensing VLMs and the first to weaponize thermal-airflow turbulence as the perturbation prior. A lightweight generator synthesizes a single input-agnostic perturbation regularized toward physically plausible airflow patterns. Optimized on one surrogate CLIP model, it attains a mean zero-shot scene-classification attack success rate (ASR, the fraction of samples whose top-1 class changes) of 48.5% across five diverse CLIP backbones, far exceeding four IR-specific physical baselines (27.7--37.0%). Applied to six state-of-the-art VLMs, it cuts scene-classification accuracy by up to 38.2% relative, yet paradoxically makes some models more confident in their IR analysis, confabulating the perturbation as genuine thermal evidence such as temperature gradients and convection. Ablations show the airflow prior raises physical plausibility at no measurable cost to attack success. Together with a benchmark spanning eleven models and four tasks, these findings expose critical vulnerabilities in the rapidly expanding IR VLM ecosystem.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.