2602.04898v1 Feb 03, 2026 cs.CR

텍스트-이미지 확산 모델에 대한 의미 수준 백도어 공격

Semantic-level Backdoor Attack against Text-to-Image Diffusion Models

Tianxin Chen
Tianxin Chen
Citations: 4
h-index: 1
Wenbo Jiang
Wenbo Jiang
Citations: 634
h-index: 10
Hongqiao Chen
Hongqiao Chen
Citations: 3
h-index: 1
Zhirun Zheng
Zhirun Zheng
Citations: 4
h-index: 1
Cheng Huang
Cheng Huang
Citations: 1
h-index: 1

텍스트-이미지(T2I) 확산 모델은 강력한 생성 능력을 가지고 널리 사용되지만, 백도어 공격에 취약합니다. 기존 공격은 일반적으로 고정된 텍스트 트리거와 단일 엔티티 백도어 목표에 의존하기 때문에, 열거 기반 입력 방어 및 어텐션 일관성 검출에 매우 취약합니다. 본 연구에서는 의미 수준에서 백도어를 심는 Semantic-level Backdoor Attack (SemBD)을 제안합니다. SemBD는 트리거를 이산적인 텍스트 패턴이 아닌 연속적인 의미 영역으로 정의하여, 표현 수준에서 백도어를 심습니다. 구체적으로, SemBD는 교차 어텐션 레이어의 핵심 및 값 투영 행렬을 증류 기반 편집을 통해 수정하여, 동일한 의미 구성을 가진 다양한 프롬프트가 백도어 공격을 안정적으로 활성화하도록 합니다. 또한, SemBD는 불완전한 의미에서의 의도치 않은 활성화를 방지하기 위한 의미 정규화와, 일관성 높은 교차 어텐션 패턴을 피하기 위한 다중 엔티티 백도어 목표를 통합하여 은밀성을 더욱 강화합니다. 광범위한 실험 결과, SemBD는 100%의 공격 성공률을 달성하면서도 최첨단 입력 수준 방어에 대한 강력한 견고성을 유지함을 보여줍니다.

Original Abstract

Text-to-image (T2I) diffusion models are widely adopted for their strong generative capabilities, yet remain vulnerable to backdoor attacks. Existing attacks typically rely on fixed textual triggers and single-entity backdoor targets, making them highly susceptible to enumeration-based input defenses and attention-consistency detection. In this work, we propose Semantic-level Backdoor Attack (SemBD), which implants backdoors at the representation level by defining triggers as continuous semantic regions rather than discrete textual patterns. Concretely, SemBD injects semantic backdoors by distillation-based editing of the key and value projection matrices in cross-attention layers, enabling diverse prompts with identical semantic compositions to reliably activate the backdoor attack. To further enhance stealthiness, SemBD incorporates a semantic regularization to prevent unintended activation under incomplete semantics, as well as multi-entity backdoor targets that avoid highly consistent cross-attention patterns. Extensive experiments demonstrate that SemBD achieves a 100% attack success rate while maintaining strong robustness against state-of-the-art input-level defenses.

1 Citations
0 Influential
5 Altmetric
26.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!