QR 구조 기반의 열적 트리거를 활용한 적외선 비전-언어 모델에 대한 표적 의미 공격
QR-Structured Thermal Triggers for Targeted Semantic Attacks on Infrared Vision-Language Models
적외선 비전-언어 모델(IR-VLMs)은 열 감지 능력을 광범위한 분류, 이미지 캡셔닝 및 시각 질의 응답으로 확장합니다. 그러나 이러한 모델들의 구조화된 열적 교란에 대한 견고성 및 다중 모달 의미 정렬의 안정성은 아직 충분히 연구되지 않았습니다. 본 논문에서는 IR-VLMs를 위한 표적 의미 조작을 가능하게 하는, 은밀하고 학습 과정이 필요 없으며 블랙박스 방식으로 작동하는 QR 구조 기반 열적 트리거(QR-STT) 프레임워크를 제안합니다. QR-STT는 QR 패턴의 기능적인 영역을 유지하면서 내부 모듈을 최적화하며, 각 모듈은 '냉', '중립' 또는 '고온'의 열 상태를 가집니다. 이 프레임워크는 위치, 크기, 회전, 강도, 블러, 원형 정도를 포함한 렌더링 파라미터와 함께 모듈 토폴로지를 동시에 탐색합니다. 세 단계로 구성된 그래디언트 기반 절차는 탐욕적인 모듈 변경을 통해 혼합된 이산적 및 연속적 검색 공간을 효율적으로 처리합니다. 이 방법은 공격자가 선택한 목표에 대한 정렬을 촉진하고, 원본 클래스에 대한 증거를 억제하며, QR 구조와 시각적 유사성을 규제하는 것을 목표로 합니다. 여러 CLIP 스타일 인코더에 대한 실험 결과, QR-STT는 일관되게 이미지-텍스트 정렬을 선택된 개념으로 재지향하면서도 시각적인 은밀성을 유지합니다. 분류를 위해 최적화된 교란은 이미지 캡셔닝 및 VQA로도 전이되어 생성된 출력에서 목표와 일치하는 의미적 변화를 유발합니다. 이러한 결과는 언어 기반 적외선 인식에 대한 QR 구조 기반 열 패턴을 해석 가능한 공격 표면으로 제시하며, 다양한 작업에서의 다중 모달 공격에 대한 견고성 평가의 필요성을 강조합니다.
Infrared vision-language models (IR-VLMs) extend thermal perception to open-vocabulary classification, image captioning, and visual question answering. However, their robustness to structured thermal perturbations and the stability of cross-modal semantic alignment remain insufficiently studied. We propose QR-Structured Thermal Triggers (QR-STT), a stealthy, training-free, black-box framework for targeted semantic steering of IR-VLMs. QR-STT preserves the functional regions of a QR pattern while optimizing its internal modules, each of which is assigned a cold, neutral, or hot thermal state. The framework jointly searches module topology and rendering parameters, including position, scale, rotation, intensity, blur, and roundness. A three-stage gradient-free procedure with greedy module-flip refinement efficiently handles the mixed discrete and continuous search space. The objective promotes alignment with an attacker-selected target, suppresses source-class evidence, and regularizes QR structure and visual similarity. Experiments on multiple CLIP-style encoders show that QR-STT consistently redirects image-text alignment toward chosen concepts while maintaining visual stealth. Perturbations optimized for classification also transfer to image captioning and VQA, causing target-consistent semantic drift in generated outputs. These results identify QR-structured thermal patterns as an interpretable attack surface for language-driven infrared perception and highlight the need for robustness evaluation against structured cross-task semantic attacks.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.