2608.10933v1 Aug 11, 2026 cs.CV

SafeCA: 안전한 교차 어텐션 지역화 및 규제 - 텍스트-비디오 생성 모델의 제로샷 공격 방어

SafeCA: Safe Cross-Attention Localization and Regulation for Text-to-Video Jailbreak Defense

Rong-Cheng Tu
Rong-Cheng Tu
Citations: 205
h-index: 7
Jiaxing Huang
Jiaxing Huang
Citations: 624
h-index: 9
Siyuan Liang
Siyuan Liang
Citations: 520
h-index: 12
Dacheng Tao
Dacheng Tao
Citations: 998
h-index: 15
Yupeng Qiu
Yupeng Qiu
Citations: 0
h-index: 0
Junfeng Fang
Junfeng Fang
Citations: 26
h-index: 2

텍스트-비디오(T2V) 생성 모델은 실제 환경에서 제로샷 공격에 취약하여 유해하거나 부적절한 콘텐츠를 생성할 수 있습니다. 기존의 방어 방법은 주로 입력 필터링 또는 재구성에 의존하는데, 이는 높은 계산 지연을 초래할 뿐만 아니라 의미를 왜곡하는 경향이 있습니다. 이러한 문제를 해결하기 위해, 우리는 실험적으로 그리고 체계적으로 클린 샘플과 제로샷 공격 샘플 간의 교차 어텐션 특징 공간 차이를 분석했습니다. 그 결과, 확산 과정 동안 두 샘플 사이에 누적적인 분리 효과가 나타나며, 선형 분리도가 점진적으로 증가한다는 사실을 처음으로 밝혀냈습니다. 이러한 통찰력을 바탕으로, 우리는 안전한 교차 어텐션을 위한 지역화 및 규제 메커니즘인 SafeCA를 제안합니다. 첫째, 단일 추론 과정에서 수집된 교차 어텐션 특징을 사용하여 어텐션 안정성 분석을 통해 핵심적인 방어 영역과 값을 식별합니다. 둘째, SafeCA는 에너지 정규화를 사용한 어텐션 마스킹을 통해 이상 활성화를 완화하고, 비정상적인 의미 흐름을 재지향하기 위해 경량의 의미 공간 어댑터를 도입합니다. 셋째, 특징 이상 신호를 입력 키워드로 역전파하여 잠재적으로 악의적인 토큰을 감지하고 억제함으로써, 상용 모델에서의 방어 적용 가능성을 향상시킵니다. 실험 결과는 SafeCA가 주류 T2V 모델에서 제로샷 공격 성공률을 약 20% 감소시키고, 거의 무시할 만한 추론 오버헤드(+0.1초)를 추가하며, 좋은 텍스트-비디오 의미 일관성을 유지한다는 것을 보여줍니다. 전반적으로 SafeCA는 T2V 생성 모델에 대한 아키텍처 수준의 방어 패러다임을 제공합니다.

Original Abstract

Text-to-Video (T2V) generative models are vulnerable to jailbreak attacks in real-world deployment, leading them to produce harmful or inappropriate content. Existing defense approaches mainly rely on input filtering or reconstruction, which not only incur high computational latency but also tend to distort semantics. To address these issues, we experimentally and systematically analyze the differences between clean and jailbreak samples in the cross-attention feature space, revealing for the first time a cumulative separation effect and a progressively increasing trend of linear separability between the two during the diffusion process. Based on this insight, we propose SafeCA, a feature-level defense mechanism for safe cross-attention localization and regularization. Firstly, we identify key defensive regions and values through attention stability analysis using cross-attention features collected from clean prompts within a single inference. Secondly, SafeCA mitigates anomalous activations via attention masking with energy normalization and introduces a lightweight semantic-space adapter to redirect abnormal semantic flows. Furthermore, we detect and suppress potentially malicious tokens by back-propagating feature anomaly signals to the input cue words, thereby enhancing the deployability of the defense in commercial models. Experimental results show that SafeCA reduces the jailbreak success rate by about 20% on mainstream T2V models, adds almost no inference overhead (+0.1s), and maintains good text-video semantic consistency. Overall, SafeCA provides an architecture-level, deployable protection paradigm for T2V generation models.

0 Citations
0 Influential
7.5 Altmetric
37.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!