2607.29445v1 Jul 31, 2026 cs.CV

QR 구조 기반의 열적 트리거를 활용한 적외선 비전-언어 모델에 대한 표적 의미 공격

QR-Structured Thermal Triggers for Targeted Semantic Attacks on Infrared Vision-Language Models

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Xiang Chen
Xiang Chen
Citations: 122
h-index: 4
Jiujiang Guo
Jiujiang Guo
Citations: 73
h-index: 5
Chengyin Hu
Chengyin Hu
Citations: 6
h-index: 1
Jiahuan Long
Jiahuan Long
Citations: 37
h-index: 4
Chao Li
Chao Li
Citations: 1
h-index: 1
Benqi Zhang
Benqi Zhang
Citations: 1
h-index: 1
Ang Li
Ang Li
Citations: 0
h-index: 0

적외선 비전-언어 모델(IR-VLMs)은 열 감지 능력을 광범위한 분류, 이미지 캡셔닝 및 시각 질의 응답으로 확장합니다. 그러나 이러한 모델들의 구조화된 열적 교란에 대한 견고성 및 다중 모달 의미 정렬의 안정성은 아직 충분히 연구되지 않았습니다. 본 논문에서는 IR-VLMs를 위한 표적 의미 조작을 가능하게 하는, 은밀하고 학습 과정이 필요 없으며 블랙박스 방식으로 작동하는 QR 구조 기반 열적 트리거(QR-STT) 프레임워크를 제안합니다. QR-STT는 QR 패턴의 기능적인 영역을 유지하면서 내부 모듈을 최적화하며, 각 모듈은 '냉', '중립' 또는 '고온'의 열 상태를 가집니다. 이 프레임워크는 위치, 크기, 회전, 강도, 블러, 원형 정도를 포함한 렌더링 파라미터와 함께 모듈 토폴로지를 동시에 탐색합니다. 세 단계로 구성된 그래디언트 기반 절차는 탐욕적인 모듈 변경을 통해 혼합된 이산적 및 연속적 검색 공간을 효율적으로 처리합니다. 이 방법은 공격자가 선택한 목표에 대한 정렬을 촉진하고, 원본 클래스에 대한 증거를 억제하며, QR 구조와 시각적 유사성을 규제하는 것을 목표로 합니다. 여러 CLIP 스타일 인코더에 대한 실험 결과, QR-STT는 일관되게 이미지-텍스트 정렬을 선택된 개념으로 재지향하면서도 시각적인 은밀성을 유지합니다. 분류를 위해 최적화된 교란은 이미지 캡셔닝 및 VQA로도 전이되어 생성된 출력에서 목표와 일치하는 의미적 변화를 유발합니다. 이러한 결과는 언어 기반 적외선 인식에 대한 QR 구조 기반 열 패턴을 해석 가능한 공격 표면으로 제시하며, 다양한 작업에서의 다중 모달 공격에 대한 견고성 평가의 필요성을 강조합니다.

Original Abstract

Infrared vision-language models (IR-VLMs) extend thermal perception to open-vocabulary classification, image captioning, and visual question answering. However, their robustness to structured thermal perturbations and the stability of cross-modal semantic alignment remain insufficiently studied. We propose QR-Structured Thermal Triggers (QR-STT), a stealthy, training-free, black-box framework for targeted semantic steering of IR-VLMs. QR-STT preserves the functional regions of a QR pattern while optimizing its internal modules, each of which is assigned a cold, neutral, or hot thermal state. The framework jointly searches module topology and rendering parameters, including position, scale, rotation, intensity, blur, and roundness. A three-stage gradient-free procedure with greedy module-flip refinement efficiently handles the mixed discrete and continuous search space. The objective promotes alignment with an attacker-selected target, suppresses source-class evidence, and regularizes QR structure and visual similarity. Experiments on multiple CLIP-style encoders show that QR-STT consistently redirects image-text alignment toward chosen concepts while maintaining visual stealth. Perturbations optimized for classification also transfer to image captioning and VQA, causing target-consistent semantic drift in generated outputs. These results identify QR-structured thermal patterns as an interpretable attack surface for language-driven infrared perception and highlight the need for robustness evaluation against structured cross-task semantic attacks.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!