디퓨전 모델의 언러닝에서 함께 연관되어 남아야 하는 개념(CARE) 연구
Co-occurring associated retained concepts in Diffusion Unlearning
언러닝은 디퓨전 모델에서 유해한 콘텐츠 생성을 완화하는 핵심 기술로 부상했습니다. 그러나 기존 방법들은 종종 목표 개념뿐만 아니라, 무해하게 동반되는 개념까지 제거하는 경향이 있습니다. 그림 1에 제시된 것처럼, 노출 장면을 제거하는 과정에서 의도치 않게 '사람'이라는 개념이 억제되어 모델이 사람 이미지를 생성하지 못하게 될 수 있습니다. 우리는 이러한 바람직하지 않은 방식으로 억제되는 동반 개념들을 CARE (Co-occurring Associated REtained concepts, 함께 연관되어 남아야 하는 개념)라고 정의합니다. 그런 다음, 우리는 CARE 점수를 제안하는데, 이는 다양한 언러닝 작업에서 이러한 개념들이 얼마나 잘 유지되는지를 직접적으로 측정하는 일반적인 지표입니다. 이를 바탕으로, 우리는 목표 개념만 제거하고 CARE를 명시적으로 보호하는 프레임워크인 ReCARE (Robust erasure for CARE)를 제안합니다. ReCARE는 대상 이미지에서 추출한 무해한 동반 토큰들의 선별된 어휘인 CARE-세트를 자동으로 구성하며, 학습 과정에서 이 어휘를 활용하여 안정적인 언러닝을 가능하게 합니다. 다양한 목표 개념(노출, 반 고흐 스타일, 텐치 객체)에 대한 광범위한 실험 결과, ReCARE는 강력한 개념 제거, 전체적인 유용성 및 CARE 유지라는 측면에서 전반적으로 최첨단 성능을 달성함을 보여줍니다.
Unlearning has emerged as a key technique to mitigate harmful content generation in diffusion models. However, existing methods often remove not only the target concept, but also benign co-occurring concepts. As illustrated in Fig.1, unlearning nudity can unintentionally suppress the concept of person, preventing a model from generating images with person. We define these undesirably suppressed co-occurring concepts that must be preserved CARE (Co-occurring Associated REtained concepts). Then, we introduce the CARE score, a general metric that directly quantifies their preservation across unlearning tasks. With this foundation, we propose ReCARE (Robust erasure for CARE), a framework that explicitly safeguards CARE while erasing only the target concept. ReCARE automatically constructs the CARE-set, a curated vocabulary of benign co-occurring tokens extracted from target images, and leverages this vocabulary during training for stable unlearning. Extensive experiments across various target concepts (Nudity, Van Gogh style, and Tench object) demonstrate that ReCARE achieves overall state-of-the-art performance in balancing robust concept erasure, overall utility, and CARE preservation.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.