Mosaic: 벡터 필드 블렌딩을 통한 장면 구성 요소 기반 다중 개념 제거
Mosaic: Compositional Multi-Concept Erasure via Vector Field Blending
텍스트-이미지(T2I) 모델에서 안전하고 윤리적인 이미지 합성을 보장하기 위한 핵심 연구 분야로 개념 제거가 부상했습니다. 기존 연구들은 여러 개념에 대한 개념 제거를 탐구했지만, 일반적으로 각 이미지당 하나의 목표 개념만 있다고 가정합니다. 이는 현대의 흐름 기반 T2I 모델이 동시에 여러 개념을 가진 복잡한 장면을 생성할 수 있게 되면서 점점 더 명확해지는 한계입니다. 이러한 격차를 해결하기 위해, 우리는 단일 장면 내에서 여러 목표 개념을 동시에 제거하는 새로운 작업인 장면 구성 요소 기반 다중 개념 제거를 소개합니다. 우리는 장면 구성 요소 기반 다중 개념 제거를 평가하기 위한 벤치마크인 CoME-Bench를 제안하며, 이는 동일 범주 및 서로 다른 범주의 시나리오를 모두 포함합니다. 또한, 흐름 기반 T2I 모델에서 다중 개념 제거를 위한 새로운 프레임워크인 Mosaic을 제안합니다. Mosaic은 벡터 필드 내의 목표 개념의 공간적 근접성을 활용하여 동적으로 개념별 마스크를 구성하고 추가적인 최적화 없이 선택적으로 블렌딩합니다. 광범위한 실험 결과는 Mosaic이 복잡한 장면 구성 요소가 있는 장면에서 여러 목표 개념을 효과적으로 제거하면서 동시에 관련 없는 배경 정보를 유지한다는 것을 보여줍니다.
Concept erasure has emerged as a key research direction for ensuring safe and ethical image synthesis in Text-to-Image (T2I) models. While existing studies have explored concept erasure across multiple concepts, they typically assume only a single target concept per image, a limitation increasingly exposed by modern flow-based T2I models, which can generate complex scenes with multiple concepts simultaneously. To address this gap, we introduce compositional multi-concept erasure, a new task that aims to simultaneously remove multiple target concepts within a single scene. We propose CoME-Bench, a benchmark for evaluating compositional multi-concept erasure, which covers both intra- and cross-category scenarios. We further propose Mosaic, a novel framework for multi-concept erasure in flow-based T2I models, which exploits the spatial locality of target concepts in the vector field by dynamically constructing concept-specific masks and selectively blending them without additional optimization. Extensive experiments demonstrate that Mosaic effectively removes multiple target concepts in complex compositional scenes while preserving non-target contexts.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.