개념 계층 구조를 갖는 확산 모델에서의 대규모 개념 제거
Mass Concept Erasure in Diffusion Models with Concept Hierarchy
확산 모델의 성공은 안전하지 않거나 유해한 콘텐츠 생성에 대한 우려를 불러일으켰으며, 이는 특정 개념을 억제하면서 일반적인 생성 능력을 유지하는 개념 제거 접근 방식의 개발로 이어졌습니다. 그러나 제거되는 개념의 수가 증가함에 따라 이러한 방법은 종종 비효율적이고 효과가 떨어지는데, 이는 각 개념마다 별도의 미세 조정 파라미터 세트가 필요하며 전체적인 생성 품질을 저하시킬 수 있기 때문입니다. 본 연구에서는 제거되는 개념을 부모-자식 구조로 구성하는 슈퍼타입-서브타입 개념 계층 구조를 제안합니다. 각 제거되는 개념은 자식 노드로 처리되며, 의미적으로 관련된 개념(예: 앵무새와 백독수리)은 공유된 부모 노드, 즉 슈퍼타입 개념(예: 새) 아래에 그룹화됩니다. 개별적으로 개념을 제거하는 대신, 우리는 의미적으로 유사한 개념을 그룹화하고 단일 세트의 학습 가능한 파라미터를 공유하여 공동으로 제거하는 효과적이고 효율적인 그룹 기반 억제 방법을 도입합니다. 제거 단계 동안 표준 확산 정규화가 적용되어 마스크되지 않은 영역에서 디노이징 프로세스를 유지합니다. 의미적으로 관련된 서브타입의 과도한 제거로 인해 발생하는 슈퍼타입 생성 성능 저하를 완화하기 위해, 우리는 슈퍼타입 개념 정보를 고정된 다운 프로젝션 행렬에 인코딩하고 제거 과정에서 업 프로젝션 행렬만 업데이트하는 새로운 방법인 Supertype-Preserving Low-Rank Adaptation (SuPLoRA)을 제안합니다. 이론적 분석은 SuPLoRA가 생성 성능 저하를 완화하는 데 효과적임을 보여줍니다. 우리는 유명인, 객체, 그리고 성적인 콘텐츠를 포함하여 다양한 도메인의 개념을 동시에 제거해야 하는 더욱 어려운 벤치마크를 구축했습니다.
The success of diffusion models has raised concerns about the generation of unsafe or harmful content, prompting concept erasure approaches that fine-tune modules to suppress specific concepts while preserving general generative capabilities. However, as the number of erased concepts grows, these methods often become inefficient and ineffective, since each concept requires a separate set of fine-tuned parameters and may degrade the overall generation quality. In this work, we propose a supertype-subtype concept hierarchy that organizes erased concepts into a parent-child structure. Each erased concept is treated as a child node, and semantically related concepts (e.g., macaw, and bald eagle) are grouped under a shared parent node, referred to as a supertype concept (e.g., bird). Rather than erasing concepts individually, we introduce an effective and efficient group-wise suppression method, where semantically similar concepts are grouped and erased jointly by sharing a single set of learnable parameters. During the erasure phase, standard diffusion regularization is applied to preserve denoising process in unmasked regions. To mitigate the degradation of supertype generation caused by excessive erasure of semantically related subtypes, we propose a novel method called Supertype-Preserving Low-Rank Adaptation (SuPLoRA), which encodes the supertype concept information in the frozen down-projection matrix and updates only the up-projection matrix during erasure. Theoretical analysis demonstrates the effectiveness of SuPLoRA in mitigating generation performance degradation. We construct a more challenging benchmark that requires simultaneous erasure of concepts across diverse domains, including celebrities, objects, and pornographic content.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.