DiffSafeMerge: 디퓨전 모델 병합 시 백도어 상속 방지
DiffSafeMerge: Mitigating Backdoor Inheritance in Diffusion Model Merging
조건부 디퓨전 체크포인트 병합은 안전한 소스를 가정하지만, 손상된 공개 체크포인트는 정상적인 이미지 생성 결과를 보이면서 잠재적인 백도어를 전달할 수 있습니다. 공격 대상, 트리거 또는 손상된 소스를 알지 못하는 경우 완화가 어렵고, 광범위한 정제 과정은 이미지 품질을 저하시킬 수 있습니다. 본 논문에서는 작은 규모의 레이블이 없는 클린 데이터셋과 고정적인, 공격에 독립적인 스트레스 프로브를 사용하여 소스 블록을 평가하고, 의심스러운 부분을 신뢰할 수 있는 참조 기준으로 축소하며, 클린 노이즈 제거 손실 예산 내에서 감쇠 값을 선택하는 DiffSafeMerge (DSM) 기법을 제안합니다. 우리는 네 가지 공격 시나리오, 두 개의 데이터셋, 그리고 21가지의 다양한 조건에 대한 실험을 수행했습니다. DSM은 14개의 소스 케이스 중 10개에서 이미 최악의 대상 ASR이 0%인 상태였으며, DSM은 이러한 결과를 유지하고 나머지 네 가지 경우에서도 세 번의 시도 동안 단 하나의 대상 매칭도 기록하지 않았습니다 (세 가지 경우는 기본 성능이 48~100%였습니다). 두 데이터셋 모두에서 최악의 대상 ASR이 0%인 방법들 중에서, DSM은 매칭된 시드-0 비교에서 가장 낮은 평균 FID 값을 보였습니다.
Unconditional diffusion checkpoint merging assumes benign sources, yet a compromised public checkpoint can transfer a dormant backdoor while clean generation appears normal. Mitigation is difficult without knowing the compromised source, trigger, or target, and broad sanitization may degrade image quality. We introduce DiffSafeMerge (DSM), which uses a small unlabeled clean set and fixed, attack-agnostic stress probes to score source blocks, shrink suspicious contributions toward a trusted reference, and select attenuation under a clean denoising-loss budget. We evaluate four attacks, two datasets, and 21 target conditions. Intended merging already has zero worst-target ASR in 10 of 14 source cases; DSM preserves these outcomes and records no target match in the remaining four over three seeds, including three with baseline ASR of 48--100\%. Among methods with zero worst-target ASR on both datasets, DSM obtains the lowest case-averaged FID in the matched seed-0 comparison.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.