다양한 디퓨전 샘플링을 위한 매니폴드 제약 기반 노이즈 최적화
Manifold-Constrained Noise Optimization for Diverse Diffusion Sampling
몇 단계로 압축된 디퓨전 모델은 고품질 이미지를 빠르게 생성하지만, 종종 프롬프트별 다양성을 잃어 랜덤 시드에 관계없이 거의 동일한 샘플을 생성합니다. 추론 시간에 초기 노이즈를 최적화하는 것은 이러한 다양성을 회복하는 매력적인 방법이지만, 기존 방법은 가우시안 사전 분포의 기하학적 구조와 모델의 노이즈 주파수 민감도를 고려하지 않고 제약 없는 유클리드 공간에서 초기 노이즈를 직접 업데이트합니다. 따라서 생성 품질을 유지하기 위해 보조 품질 관리 목표를 추가하고 계산 비용과 가중치 하이퍼파라미터를 사용하며, 동시에 성능 저하를 방지하기 위해 보수적인 업데이트가 필요합니다. 본 연구에서는 학습이 필요 없는 MoNO라는 방법을 제안합니다. MoNO는 저차원이며 품질을 안정화하는 노이즈 매니폴드에서 매니폴드 제약 기반 노이즈 최적화를 수행합니다. MoNO는 예측된 시각적 특징이 이전 생성 결과와 상호 보완되도록 각 새로운 초기 노이즈를 순차적으로 최적화하며, 아핀 저주파 구체에서의 리만 업데이트는 사전 확률을 유지하고 불안정한 고주파 성분을 제거합니다. 이를 통해 큰 기하 곡선 경로 탐색이 가능하며, 보조 품질 관리 목표의 필요성을 없애고 기존 노이즈 최적화 방법보다 훨씬 적은 반복 횟수로 수렴됩니다. 여러 압축된 텍스트-이미지 디퓨전 모델에 대한 실험 결과, MoNO는 이미지 품질을 유지하면서 프롬프트별 다양성을 꾸준히 향상시키는 것으로 나타났습니다.
Few-step distilled diffusion models generate high-quality images quickly, but often lose per-prompt diversity, producing near-identical samples across random seeds. Optimizing the initial noise at inference time offers an appealing way to recover this diversity, yet existing methods directly update the initial noise in an unconstrained Euclidean space, ignoring both the geometry of the Gaussian prior and the model's sensitivity to noise frequencies. They therefore introduce auxiliary quality-control objectives to maintain generation fidelity, adding compute and weighting hyperparameters while still requiring conservative updates to prevent degradation. In this work, we propose MoNO, a training-free method that performs Manifold-constrained Noise Optimization on a low-dimensional, quality-stabilizing noise manifold. MoNO sequentially optimizes each new initial noise so that its predicted visual feature complements previous generations, while Riemannian updates on an affine low-frequency sphere preserve prior likelihood and fix unstable high-frequency components by construction. This enables large geodesic steps, removes the need for auxiliary quality-control objectives, and converges in far fewer iterations than prior noise-optimization methods. Experiments with multiple distilled text-to-image diffusion models show that MoNO consistently improves per-prompt diversity while maintaining image quality.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.