2608.09244v1 Aug 10, 2026 cs.CV

결합된 잠재 노이즈 가이드 기반의 루프 내 모델 적응을 통한 고품질 주제 중심 텍스트-이미지 생성

In-Loop Model Adaptation with Coupled Latent-Noise Guidance for High-Fidelity Subject-Driven Text-to-Image Generation

Weiming Chen
Weiming Chen
Citations: 0
h-index: 0
Siyi Liu
Siyi Liu
Citations: 95
h-index: 3
Yushun Tang
Yushun Tang
Citations: 162
h-index: 8
Yi Zhang
Yi Zhang
Citations: 106
h-index: 5
Feng Wu
Feng Wu
Citations: 2,619
h-index: 22
Zhihai He
Zhihai He
Citations: 279
h-index: 9

텍스트-이미지 확산 모델은 주어진 텍스트 프롬프트로부터 고품질 이미지를 생성하는 데 놀라운 성공을 거두었습니다. 주제 중심 생성은 주어진 참조 이미지에 나타난 객체의 특징을 모방하면서 다양한 시각적 맥락을 텍스트 프롬프트를 통해 지정하여 맞춤형 이미지를 합성하는 것을 목표로 합니다. 핵심적인 과제는 참조 이미지가 변경될 때, 확산 모델이 서로 다른 시각적 맥락에 효율적으로 적응하고 동시에 객체의 동일성을 일관되게 유지해야 한다는 것입니다. 기존 방법은 모델을 대규모의 도메인 특화 데이터셋으로 학습시키거나, 실제 이미지 생성이 시작되기 전에 참조 이미지를 사용하여 수백 번의 반복 훈련(fine-tuning)을 수행합니다. 본 연구에서는 extit{루프 내 모델 적응 (In-Loop Model Adaptation, IMA)}이라는 새로운 방법을 제안합니다. 이 방법은 실제 이미지 생성 과정 중 각 단계에서 핵심 확산 모델을 조정하며, 이미지 생성 프로세스 전에 참조 이미지를 사용하여 학습하지 않습니다. 이를 위해, 우리는 참조 이미지를 일련의 잠재 변수로 매핑하는 DDIM 역전환 체인과 텍스트 프롬프트만을 사용하여 이미지를 생성하는 텍스트-이미지 생성 체인을 구축합니다. 그런 다음, 각 생성 단계에서 확산 모델과 이 두 체인 간의 잠재-노이즈 차이를 특징짓기 위해 마스크된 잠재 일관성 손실 및 노이즈 정규화 손실을 도입합니다. 이러한 결합된 잠재-노이즈 손실은 루프 내 모델 적응을 안내하여 참조 이미지에 의해 지정된 객체 동일성을 유지하면서 텍스트 프롬프트와의 정확한 정렬을 유지하도록 하여 고품질의 텍스트-이미지 생성을 가능하게 합니다. 광범위한 실험 결과, 제안하는 IMA 방법이 주제 중심 텍스트-이미지 생성 성능을 크게 향상시키는 것을 보여줍니다.

Original Abstract

Text-to-image diffusion models have achieved remarkable success in generating high-quality images from a given text prompt. Subject-driven generation aims to synthesize customized images to mimic the appearance of subjects in given reference images within different visual contexts specified by the text prompts. The central challenge here is that, when the reference image changes, the diffusion model cannot efficiently adapt to different visual contexts while consistently maintaining the subject identity. Existing methods either train the model with a large domain-specific dataset or fine-tune the model using the reference image for hundreds of iterations before actual image generation. In this work, we explore a new approach, called \textit{In-Loop Model Adaptation} (IMA), which adapts the core diffusion model at each generation step during the actual process of image generation, without being trained on the reference image before the generation process. To this end, we establish a DDIM inversion chain that maps the reference image to a sequence of latent, as well as a text-to-image generation chain which generates the image from the text prompt only. We then introduce a masked latent consistency loss and a noise regularization loss to characterize the latent-noise difference between the diffusion model and these two chains at each generation step. This coupled latent-noise loss is used to guide the in-loop model adaptation to preserve the subject identity specified by the reference image while maintaining accurate alignment with the text prompt, resulting in high-fidelity text-to-image generation. Our extensive experiments demonstrate that our proposed IMA method significantly improves the performance of subject-driven text-to-image generation.

0 Citations
0 Influential
11 Altmetric
55.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!