2604.16481v1 Apr 12, 2026 cs.CV

수천 개의 개념 삭제: 텍스트-이미지 확산 모델을 위한 확장 가능하고 실용적인 개념 삭제 방법

Erasing Thousands of Concepts: Towards Scalable and Practical Concept Erasure for Text-to-Image Diffusion Models

Se Young Chun
Se Young Chun
Citations: 43
h-index: 4
H. Seo
H. Seo
Citations: 113
h-index: 5
Byung Hyun Lee
Byung Hyun Lee
Citations: 173
h-index: 5
Jaehyun Cho
Jaehyun Cho
Citations: 2
h-index: 1
Sungjin Lim
Sungjin Lim
Citations: 47
h-index: 2

대규모 텍스트-이미지(T2I) 확산 모델은 놀라운 시각적 충실도를 제공하지만, 저작권 콘텐츠와 같은 바람직하지 않은 콘텐츠를 재현할 수 있는 능력으로 인해 안전 문제를 야기합니다. 개념 삭제는 이러한 문제를 완화하기 위한 전략으로 부상했지만, 기존 방법은 확장성, 정확성 및 견고성 간의 균형을 맞추는 데 어려움을 겪으며, 결과적으로 수백 개의 개념만 삭제하는 데 제한됩니다. 이러한 제한 사항을 해결하기 위해, 우리는 수천 개의 개념을 삭제하면서도 생성 품질을 유지할 수 있는 확장 가능한 프레임워크인 Erasing Thousands of Concepts (ETC)를 제안합니다. 우리의 방법은 먼저 스튜어트 t-분포 혼합 모델(tMM)을 사용하여 저차원 개념 분포를 모델링합니다. 이를 통해 affine 최적 수송을 활용하여 대상 개념을 정확하게 삭제하고, 동시에 사전 정의된 기준 개념 없이 대상 개념 분포의 경계를 고정하여 다른 개념은 보존합니다. 그런 다음, Mixture-of-Experts (MoE) 기반 모듈인 MoEraser를 훈련시켜 대상 임베딩은 제거하고 기준 임베딩은 유지합니다. 텍스트 임베딩 프로젝터에 노이즈를 주입하고 MoEraser를 미세 조정함으로써, 우리의 프레임워크는 모듈 제거와 같은 white-box 공격에 대한 견고성을 달성합니다. 다양한 도메인과 확산 모델에 걸쳐 2,000개 이상의 개념에 대한 광범위한 실험을 통해, 우리의 프레임워크가 대규모 개념 삭제에서 최첨단 수준의 확장성과 정확성을 제공한다는 것을 입증했습니다.

Original Abstract

Large-scale text-to-image (T2I) diffusion models deliver remarkable visual fidelity but pose safety risks due to their capacity to reproduce undesirable content, such as copyrighted ones. Concept erasure has emerged as a mitigation strategy, yet existing approaches struggle to balance scalability, precision, and robustness, which restricts their applicability to erasing only a few hundred concepts. To address these limitations, we present Erasing Thousands of Concepts (ETC), a scalable framework capable of erasing thousands of concepts while preserving generation quality. Our method first models low-rank concept distributions via a Student's t-distribution Mixture Model (tMM). It enables pin-point erasure of target concepts via affine optimal transport while preserving others by anchoring the boundaries of target concept distributions without pre-defined anchor concepts. We then train a Mixture-of-Experts (MoE)-based module, termed MoEraser, which removes target embeddings while preserving the anchor embeddings. By injecting noise into the text embedding projector and fine-tuning MoEraser for recovery, our framework achieves robustness to white-box attack such as module removal. Extensive experiments on over 2,000 concepts across heterogeneous domains and diffusion models demerate state-of-the-art scalability and precision in large-scale concept erasure.

1 Citations
0 Influential
2.5 Altmetric
13.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!