2606.31699v1 Jun 30, 2026 cs.CV

디퓨전 모델에서 희소 오토인코더를 활용한 지식 삭제: 보고는 하지만 건드리지 마세요

Look But Don't Touch with Sparse Autoencoders for Unlearning in Diffusion Models

Riccardo Renzulli
Riccardo Renzulli
Citations: 101
h-index: 5
Marco Grangetto
Marco Grangetto
Citations: 140
h-index: 7
S. Alaniz
S. Alaniz
Citations: 620
h-index: 14
E. Cassano
E. Cassano
Citations: 9
h-index: 2
Rayyan Ahmed
Rayyan Ahmed
Citations: 4
h-index: 1

희소 오토인코더(SAE)는 최근 개념 수준의 조작을 위한 해석 가능한 도구로 제안되었으며, 이는 분리된 특징들이 제어 가능한 개입점으로 작용할 수 있다는 가정하에 이루어졌습니다. 본 연구에서는 디퓨전 모델에서의 객체 제거 및 방향 조절이라는 맥락에서 이러한 가정을 체계적으로 평가합니다. SAE가 디퓨전 모델 활성화 내의 의미적 개념을 안정적으로 감지하고 위치시키는 데 효과적이지만, 잠재 공간에 직접적인 개입을 시도하면 종종 데이터 분포 외부의 활성화를 유발하여 심각한 시각적 왜곡이 발생한다는 것을 확인했습니다. 감지와 개입을 분리하기 위해, 우리는 SAE 활성화를 순전히 의미 감지기로 활용하여 대상 객체를 포함하는 이미지 영역을 식별하고, 해당 패치 임베딩을 객체를 포함하지 않는 임베딩으로 대체합니다. 이러한 감지 기반의 치환은 디퓨전 모델의 활성화 통계치를 유지하며, 잠재 공간 조작보다 훨씬 깨끗한 제거 결과를 제공합니다. 우리의 연구 결과는 디퓨전 모델에서 개념 감지와 개념 개입 사이에 근본적인 간극이 존재함을 보여줍니다. 즉, 단일 의미 또는 희소 특징은 방향 조절을 위한 적합한 제어 장치가 아닙니다. 이러한 결과는 SAE를 생성 모델 분석을 위한 강력한 해석 도구로 자리매김시키지만, 직접적인 조작(예: 지식 삭제)에 사용할 때 중요한 한계점을 강조합니다.

Original Abstract

Sparse autoencoders (SAEs) have recently been proposed as interpretable tools for concept-level manipulation, under the assumption that isolated features can serve as controllable intervention points. In this work, we systematically evaluate this assumption in the context of object erasure and steering in diffusion models. We show that while SAEs reliably detect and localize semantic concepts within diffusion model activations, direct intervention in their latent space frequently induces out-of-distribution activations, resulting in severe visual artifacts. To disentangle detection from intervention, we use SAE activations purely as semantic detectors to identify image regions containing the target object, and replace those patch embeddings with the ones that do not contain it. This detection-based replacement preserves the diffusion model's activation statistics and produces significantly cleaner erasure results than latent steering. Our findings reveal a fundamental gap between concept detection and concept intervention in diffusion models: monosemantic or sparse features are not inherently suitable as control knobs for steering. These results position SAEs as powerful interpretability tools for analyzing generative models, but highlight important limitations when used for direct manipulation, such as unlearning.

0 Citations
0 Influential
7 Altmetric
35.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!