텍스트-이미지 생성의 충실도 향상: 추론 시 안정화 및 제어
Anchoring and Steering Diffusion: Enhancing the Faithfulness of Text-to-Image Generation at Inference Time
텍스트-이미지 확산 모델은 뛰어난 시각적 품질을 달성하지만, 복잡한 구성 프롬프트에 대한 정확성을 유지하는 데 어려움을 겪는 경우가 많습니다. 효과적인 전략은 확산 모델의 추론 프로세스를 개선하여 사전 학습된 지식을 활용하고 불일치를 해결하는 것입니다. 기존의 학습이 필요 없는 방법은 크게 두 가지 범주로 나눌 수 있습니다. 첫 번째 범주는 무작위로 샘플링된 초기 노이즈를 개선하는 데 중점을 두며, 이는 비용이 많이 드는 노이즈 풀 검색을 수행하거나 신뢰할 수 있는 의미 주입을 보장하지 않고 샘플링된 노이즈를 조작하는 것을 포함합니다. 두 번째 범주에서는 디노이징 경로를 개선하는 데 중점을 두지만, 의미 오류를 적시에 진단하고 수정하기 위한 명시적인 메커니즘이 부족합니다. 본 연구에서는 초기화와 디노이징 경로 모두에 대한 세밀한 제어를 수행하는 학습이 필요 없는 프레임워크인 AnchorSteer를 제안합니다. AnchorSteer는 두 가지 상호 보완적인 구성 요소로 구성됩니다: Semantic Anchoring은 CLIP 기반 사전 추출 및 새로운 Latent-Prior Score Distillation Sampling (LP-SDS) 목표를 통해 유용한 정보를 제공하지 않는 가우시안 노이즈를 텍스트에 맞춰진 초기화 값으로 대체합니다. 특히, LP-SDS는 CLIP 시각적 사전을 확산 모델의 지식 분포로 전달하여 CLIP 기반 사전과 확산 기반 사전 간의 도메인 격차를 줄입니다. Reflective Steering는 Think--Erase--Retouch 루프를 통해 수동적인 디노이징을 능동적인 방식으로 변환하여 생성 과정 중에 자체 수정이 가능합니다. 이는 VLM 기반 진단을 활용하여 의미적 편차를 감지하고, 오류가 있는 콘텐츠를 억제하고 누락된 속성을 복구하기 위해 대상 레이턴트 공간 개선을 수행합니다. GenEval 및 T2I-CompBench++에 대한 광범위한 실험 결과, AnchorSteer는 텍스트-이미지 정렬 측면에서 기존 방법보다 우수한 성능을 보이며 높은 시각적 품질을 유지하는 것으로 나타났습니다.
While text-to-image diffusion models achieve impressive visual quality, they frequently struggle to maintain precise alignment with complex compositional prompts. An effective strategy is to improve the inference process of diffusion models, thereby better leveraging their pretrained priors to address misalignment. Existing training-free methods can be divided into two categories. The first category focuses on improving the randomly sampled initial noise, either performing costly search over noise pools or manipulating sampled noise without ensuring reliable semantic injection. The second category focuses on improving the denoising trajectory, lacking explicit mechanisms to timely diagnose and correct semantic errors. we propose \textbf{AnchorSteer}, a training-free framework that exerts fine-grained control over \textbf{both initialization} and \textbf{the denoising trajectory}. AnchorSteer consists of two synergistic components: \textbf{Semantic Anchoring} replaces uninformative Gaussian noise with text-aligned initializations via CLIP-based prior extraction and a novel Latent-Prior Score Distillation Sampling (LP-SDS) objective. Specifically, LP-SDS distills CLIP visual priors into the knowledge distribution of diffusion models, mitigating the domain gap between CLIP-based priors and diffusion-based priors. \textbf{Reflective Steering} transforms passive denoising with an active Think--Erase--Retouch loop that enables mid-generation self-correction. It leverages VLM-based diagnosis to detect semantic deviations and performs targeted latent refinement to suppress erroneous content and recover missing attributes. Extensive experiments on GenEval and T2I-CompBench++ demonstrate that AnchorSteer consistently outperforms existing baselines in text--image alignment while preserving high visual quality.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.