2607.06432v1 Jul 07, 2026 cs.LG

TILDE: 기울어진 분포 기반의 개념 삭제를 통한 개념 제거

TILDE: TILt-based Distributional Erasure for Concept Unlearning

Yuhta Takida
Yuhta Takida
Citations: 1,365
h-index: 16
N. Murata
N. Murata
Citations: 1,409
h-index: 17
Yuki Mitsufuji
Yuki Mitsufuji
Citations: 1,716
h-index: 16
Naveen George
Naveen George
Citations: 27
h-index: 1
Konda Reddy Mopuri
Konda Reddy Mopuri
Indian Institute of Technology Guwahati
Citations: 3,148
h-index: 18

텍스트-이미지 확산 모델에서 개념 제거는 안전하고 실용적인 배포에 매우 중요합니다. 개인 정보 보호 문제, 저작권 분쟁, 상표 제한 및 안전 규제가 증가함에 따라, 배포된 시스템은 훈련 후 원치 않는 개념을 효과적으로 제거할 수 있어야 합니다. 기존 방법들은 일반적으로 목표 개념을 잘 제거하지만, 실제 개념 제거에는 동등하게 중요한 특성이 필요합니다. 즉, 제거된 모델은 품질, 다양성 및 의미적 범위를 유지해야 하며, 이는 양호한 이미지 생성을 가능하게 해야 합니다. 이상적인 모델은 원치 않는 데이터를 사용하지 않고 처음부터 훈련된 모델이며, 이 모델은 기존 모델과 유사한 성능을 보여야 합니다. 그러나 일반적인 제거 목적 함수는 제거 후 어떤 분포가 이러한 기준 분포를 근사해야 하는지 명시적으로 정의하지 않기 때문에, 유지 능력은 업데이트 규칙의 부산물로 간주됩니다. 본 논문에서는 TILDE라는 새로운 방법을 제안합니다. TILDE는 기울어진 분포 기반의 개념 삭제 방법으로, 개념 제거를 분포 정렬 문제로 공식화합니다. 즉, 목표는 사전 훈련된 모델에서 벗어난 최소 편차를 갖는 조건부 분포이며, 이는 망각 제약 조건을 고려합니다. 이러한 에너지 기울기를 이용한 방식으로, 원치 않는 이미지 표현을 억제하는 동시에 각 프롬프트에 대한 양호한 이미지 생성 가능성을 유지합니다. 우리는 잔여 $ abla$-GFlowNet 훈련 방식을 사용하여 사전 훈련된 확산 모델에 상대적인 망각 에너지로 인해 발생하는 점수 보정을 학습합니다. TILDE는 다양한 객체, 예술 스타일 및 캐릭터에 대해 기존 방법보다 우수한 성능을 보이면서도, 강력한 망각 능력을 제공하고 유지 능력과 분포 충실도를 향상시킵니다.

Original Abstract

Concept unlearning in text-to-image diffusion models is critical for safe and practical deployment: with rising privacy concerns, copyright disputes, trademark constraints, and safety regulations, deployed systems must be able to suppress unwanted concepts after training. Existing methods often remove the target concept effectively, but practical unlearning also requires an equally fundamental property: the unlearned model should retain quality, diversity, and semantic coverage on benign generation. The gold standard is a retain-only model trained from scratch without the unwanted data. However, common erasure objectives do not specify which post-unlearning distribution should approximate this reference, leaving retention as an implicit consequence of the update rule. We propose TILDE, TILt-based Distributional Erasure, which formulates concept unlearning as a distributional alignment problem: the desired target is the minimum-deviation conditional distribution from the pretrained model under a forgetting constraint. This energy-tilted, anchor-free target suppresses concept-expressing images while preserving benign relative mass for each prompt. We instantiate this principle with residual $\nabla$-GFlowNet training, which learns the score correction induced by the forget energy relative to the pretrained diffusion model. Across objects, artistic styles, and characters, TILDE achieves strong forgetting while improving retention and distributional fidelity over prior baselines.

0 Citations
0 Influential
9 Altmetric
45.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!