표현 정렬 투영기를 사용한 확산 모델의 학습-불필요 표현 가이드
Training-Free Representation Guidance for Diffusion Models with a Representation Alignment Projector
최근 생성 모델 분야의 발전은 확산 기반 프레임워크를 통해 고품질의 시각적 합성 기능을 가능하게 했으며, 이는 제어 가능한 샘플링과 대규모 학습을 지원합니다. 분류기-없는 방식 및 대표 가이드와 같은 추론 시간 가이드 방법은 샘플링 동역학을 수정하여 의미적 정렬을 향상시키지만, 비지도 특징 표현을 충분히 활용하지 못합니다. 이러한 시각적 표현은 풍부한 의미 구조를 포함하고 있지만, 추론 시에 실제 정답 이미지(ground-truth reference images)가 없기 때문에 생성 과정에서 이러한 표현을 통합하는 데 제약이 있습니다. 본 연구에서는 확산 트랜스포머의 초기 노이즈 제거 단계에서 발생하는 의미적 편향을 밝히고, 동일한 조건 하에서도 확률성으로 인해 일관성 없는 정렬이 발생하는 현상을 분석합니다. 이 문제를 해결하기 위해, 표현 정렬 투영기를 사용하는 가이드 방식을 제안합니다. 이 방식은 투영기가 예측한 표현을 중간 샘플링 단계에 주입하여 모델 아키텍처를 변경하지 않고 효과적인 의미적 기준점을 제공합니다. SiTs 및 REPAs에 대한 실험 결과, 클래스 조건부 ImageNet 합성에서 상당한 성능 향상이 확인되었으며, FID 점수가 크게 감소했습니다. 예를 들어, REPA-XL/2의 경우 5.9에서 3.3으로 개선되었으며, 제안된 방법은 SiT 모델에 적용했을 때 대표 가이드보다 우수한 성능을 보였습니다. 또한, 본 접근 방식은 분류기-없는 가이드와 결합될 때 상호 보완적인 성능 향상을 가져와 의미적 일관성과 시각적 충실도를 향상시키는 것으로 나타났습니다. 이러한 결과는 표현 기반의 확산 샘플링이 의미적 보존과 이미지 일관성을 강화하는 실용적인 전략임을 입증합니다.
Recent progress in generative modeling has enabled high-quality visual synthesis with diffusion-based frameworks, supporting controllable sampling and large-scale training. Inference-time guidance methods such as classifier-free and representative guidance enhance semantic alignment by modifying sampling dynamics; however, they do not fully exploit unsupervised feature representations. Although such visual representations contain rich semantic structure, their integration during generation is constrained by the absence of ground-truth reference images at inference. This work reveals semantic drift in the early denoising stages of diffusion transformers, where stochasticity results in inconsistent alignment even under identical conditioning. To mitigate this issue, we introduce a guidance scheme using a representation alignment projector that injects representations predicted by a projector into intermediate sampling steps, providing an effective semantic anchor without modifying the model architecture. Experiments on SiTs and REPAs show notable improvements in class-conditional ImageNet synthesis, achieving substantially lower FID scores; for example, REPA-XL/2 improves from 5.9 to 3.3, and the proposed method outperforms representative guidance when applied to SiT models. The approach further yields complementary gains when combined with classifier-free guidance, demonstrating enhanced semantic coherence and visual fidelity. These results establish representation-informed diffusion sampling as a practical strategy for reinforcing semantic preservation and image consistency.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.