2607.25522v1 Jul 28, 2026 cs.CV

I2VShield: Diffusion Transformer 기반 이미지-비디오 모델에 대한 효율적인 사전 방어 프레임워크

I2VShield: An Efficient Proactive Defense Framework against DiT-based Image-to-Video Models

Wei Lu
Wei Lu
Citations: 220
h-index: 9
Yimao Guo
Yimao Guo
Citations: 2
h-index: 1
Zuomin Qu
Zuomin Qu
Citations: 52
h-index: 3

영상 생성 모델의 빠른 발전으로 인해 이미지-비디오(I2V) 모델의 오용이 증가하고 있습니다. AI가 생성한 영상을 탐지하는 데 상당한 진전이 있었지만, I2V 모델에 대한 사전 방어는 아직 연구가 부족합니다. 특히, 현재 I2V 모델에 대한 대부분의 사전 방어 방법은 그래디언트 기반 적대적 공격에 의존하며, 이는 방어자가 적대적인 예제를 생성하기 위해 상당한 메모리 자원(VRAM)을 갖춘 GPU를 필요로 합니다. 이러한 문제를 해결하기 위해, 우리는 Diffusion Transformer (DiT) 기반 I2V 모델에 특화된 생성적 적대적 공격을 기반으로 하는 개인 정보 보호 방법인 I2VShield를 제안합니다. 제안하는 방법은 크게 두 가지 구성 요소로 이루어집니다: (1) 계산 오버헤드를 줄이면서도 시각적으로 인지하기 어렵도록 적대 학습을 통합한 텍스트-적응형 이상치 생성 프레임워크; 그리고 (2) DiT 기반 I2V 모델의 고유한 취약점을 활용하여 내부 어텐션 특징의 클린 상태로부터의 편차를 최대화하는 비표적 다중 모드 어텐션 방해(MAD) 공격입니다. 광범위한 실험 결과, 제안하는 방법은 다양한 데이터셋과 주류 DiT 기반 I2V 모델에서 뛰어난 보호 성능을 달성하며, 특히 시공간적 일관성을 저해하는 동시에 계산 비용을 크게 줄이는 것을 보여줍니다.

Original Abstract

The rapid advancement of video generation models has led to the increasing misuse of image-to-video (I2V) models. Although substantial progress has been made in detecting AI-generated videos, proactive defenses against I2V models remain underexplored. In particular, current proactive defenses against I2V models predominantly rely on gradient-based adversarial attacks, which require defenders to possess GPUs with substantial memory resources (VRAM) to generate adversarial examples. To address this issue, we propose I2VShield, a privacy protection method based on generative adversarial attacks tailored to Diffusion Transformer (DiT)-based I2V models. The proposed method primarily consists of two components: (1) a text-adaptive perturbation generation framework integrating adversarial learning to mitigate computational overhead while maintaining visual imperceptibility; and (2) an untargeted Multimodal Attention Disruption (MAD) attack that exploits the inherent vulnerabilities of DiT-based I2V models, maximizing the deviation of the internal attention features from their clean states. Extensive experiments demonstrate that our approach achieves highly competitive protection performance across various datasets and mainstream DiT-based I2V models, particularly in disrupting spatiotemporal coherence, while substantially reducing computational costs.

0 Citations
0 Influential
4.5 Altmetric
22.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!