텍스트-이미지 확산 트랜스포머를 위한 강력하고 일반화된 안전 제어 방법
Robust and Generalizable Safety Steering for Text-to-Image Diffusion Transformers
확산 트랜스포머는 텍스트-이미지 생성의 핵심 기술로 자리 잡았지만, 계층적이고 크로스 모달 방식의 생성 과정은 프롬프트 수준 필터링이나 출력 수준 감지와는 근본적으로 다른 안전 제어 방식을 요구합니다. 유해한 의미 정보는 텍스트 표현에 미약하게 표현될 수 있으며, 점진적으로 시각적 잠재 공간에 연결되고, 최종적으로 렌더링 역학에 얽히게 됩니다. 따라서 특정 계층에서 안전 제어를 수행하는 것은 불안정할 수 있으며, 알려진 위험으로부터 학습된 제어 메커니즘이 새로운 위험 영역으로 쉽게 이전되지 않을 수 있습니다. 본 연구에서는 DiT의 안전 적응을 위치 정보를 고려한 희소 특징 전송으로 정의하는 안전 제어 프레임워크인 SafeDIG를 제안합니다. SafeDIG는 기능적으로 구별되는 DiT 개입 위치에 대한 희소 자동 인코더(Sparse Autoencoders, SAE)를 구축하고, 소스-타겟 위험 변화에 안정적일 것으로 예상되는 개입 지점을 우선시하기 위해 견고성 기반 사전 학습 라우팅을 사용합니다. 또한, 재사용 가능한 희소 안전 딕셔너리로 SAE 인코더를 고정하고 디코더만 타겟 도메인 활성화 공간에 적응시켜 전이 가능한 안전 특징과 도메인별 활성화 형상을 분리합니다. 추론 과정에서 SafeDIG는 Blend 및 Repel 연산을 결합하여 유해한 활성화를 전이된 안전 공간으로 이동시키거나, 유해한 희소 방향으로부터 멀어지도록 제어합니다. FLUX.1 Dev 및 Stable Diffusion 3.5 Large 데이터셋에 대한 실험 결과, SafeDIG는 타겟 도메인과 전체적인 유해 생성 비율을 지속적으로 감소시키는 동시에 소스 도메인의 안전성과 이미지 품질을 유지하는 것으로 나타났습니다.
Diffusion Transformers have become a powerful backbone for text-to-image generation, but their layered and cross-modal generation process makes safety control fundamentally different from prompt-level filtering or output-level detection. Harmful semantics may be weakly expressed in text representations, progressively bound to visual latents, and finally entangled with rendering dynamics. As a result, safety steering at a fixed layer can be unstable, and a steering mechanism learned from known risks may not transfer reliably to a shifted target risk domain. We propose SafeDIG, a safety steering framework that formulates DiT safety adaptation as position-aware sparse feature transfer. SafeDIG first constructs Sparse Autoencoders over functionally distinct DiT intervention positions and uses robustness-aware pre-training routing to prioritize intervention sites that are expected to remain stable under source-target risk shift. It then separates transferable safety features from domain-specific activation geometry by freezing the SAE encoder as a reusable sparse safety dictionary and adapting only the decoder to the target-domain activation manifold. During inference, SafeDIG combines Blend and Repel operations to steer unsafe activations toward transferred safety manifolds or away from harmful sparse directions. Experiments on FLUX.1 Dev and Stable Diffusion 3.5 Large show that SafeDIG consistently reduces target-domain and overall unsafe generation rates while preserving source-domain safety and image quality.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.