TIGA: 경로 주입 생성 공격 - 블랙박스 AIGC 탐지기를 위한 방법
TIGA: Trajectory-Injected Generative Attack against Black-box AIGC Detectors
최근의 확산 모델은 얼굴 이미지 합성에서 놀라운 수준의 현실감을 보여주며, 인공지능 생성 콘텐츠(AIGC) 포렌식 탐지기에 큰 어려움을 야기하고 있습니다. 기존의 회피 기법들은 일반적으로 이미 생성된 이미지에 노이즈를 추가하거나, 탐지기를 고려한 학습을 필요로 하며, 이는 눈에 띄는 또는 통계적인 왜곡을 초래할 수 있으며, 확산 모델이 고정되어 있어야 하고 대상 탐지기가 블랙박스 형태로만 접근 가능하다는 제약 하에서는 적용 가능성이 제한됩니다. 본 논문에서는 원본 이미지 없이, 학습 과정 없이도 탐지기를 회피하는 이미지를 단일 확산 샘플링 경로 내에서 생성하는 프레임워크인 Trajectory-Injected Generative Attack (TIGA)를 제안합니다. TIGA는 적대적인 특성이 생성 과정 중에 나타나도록 잠재된 Denoising Diffusion Implicit Model (DDIM)의 경로를 조정합니다. TIGA는 먼저 여러 개의 화이트박스 대리 탐지기에서 얻은 기울기를 집계하여 전이 가능한, 부호 정보를 포함하는 사전 지식을 형성하고, 이어서 블랙박스 대상 시스템의 응답을 추정하기 위해 대칭적인 유한 차분 쿼리를 사용한 등방성 방향 탐색을 수행합니다. 추정된 방향은 감쇠된 모멘텀으로 안정화되고, DDIM 노이즈 스케줄에 따라 주입되며, 고주파 왜곡을 줄이기 위해 주파수 영역에서 변형됩니다. 대리 탐지기와 알려지지 않은 특수 포렌식 탐지기에 대한 실험 결과, TIGA는 원본 이미지나 확산 모델 재학습 없이도 강력한 블랙박스 공격 성능, 전이성 및 높은 견고성을 달성하며, 동시에 높은 시각적 품질을 유지함을 보여줍니다.
Recent diffusion models have achieved remarkable realism in facial image synthesis, posing growing challenges to artificial intelligence-generated content (AIGC) forensic detectors.Existing evasion methods typically perturb pre-generated images or require detector-aware training, which may introduce visible or statistical artifacts and limit applicability when the diffusion model must remain frozen and the target detector is accessible only through black-box queries. We propose Trajectory-Injected Generative Attack (TIGA), a source-image-free and training free framework that generates detector-evasive images within a single diffusion sampling trajectory. TIGA steers the latent Denoising Diffusion Implicit Model (DDIM) trajectory so that adversarial properties emerge during generation rather than being added afterward. TIGA first aggregates gradients from multiple white-box surrogate detectors to form a transferable, sign-aware prior, and then performs anisotropic directional search with symmetric finite-difference queries to estimate the black-box target response. The estimated directions are stabilized by decayed momentum and injected according to the DDIM noise schedule, with frequency-domain reshaping to suppress high frequency artifacts. Experiments on surrogate and unseen specialized forensic detectors show that TIGA achieves strong blackbox attack performance, transferability, and high robustness under common post-processing operations without source images or diffusion-model retraining, while preserving high perceptual quality.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.