2605.26032v1 May 25, 2026 cs.CV

모든 크기에 적용 가능한 이미지 생성: 연속적인 초해상도를 갖는 스케일 불변 확산 모델

Everything at Every Scale: Scale-Invariant Diffusion with Continuous Super-Resolution

Jeff Gore
Jeff Gore
Citations: 75
h-index: 5
Z. Chen
Z. Chen
Citations: 4
h-index: 1
Archer Wang
Archer Wang
Citations: 6
h-index: 2
Marin Soljačić
Marin Soljačić
Citations: 1,972
h-index: 9
William T. Freeman
William T. Freeman
Citations: 1,752
h-index: 5
Congyue Deng
Congyue Deng
Citations: 1,176
h-index: 15
Zhuo Chen
Zhuo Chen
Citations: 105
h-index: 5

노이즈로부터 이미지를 생성하는 것은 이미지 생성이며, 낮은 해상도 데이터로부터 미세한 디테일을 복원하는 것은 초해상도입니다. 실질적인 차이점에도 불구하고, 이 두 가지 작업은 모두 스케일 간 정보 손실을 역전시키는 것으로 이해할 수 있습니다. 본 논문에서는 $ extbf{SKILD}$라는 $ extbf{S}$케일 불변 $ extbf{K}$-공간 $ extbf{I}$이미지 $ extbf{L}$러닝 $ extbf{D}$iffusion 모델을 소개합니다. 이 모델은 이미지 생성과 연속적인 초해상도를 단일하고 조건 없는 프레임워크 내에서 통합합니다. 자연 이미지와 중요한 물리 시스템 모두 스케일 불변성을 나타내며, 우리는 이를 활용하여 forward 과정(이미지 콘텐츠를 미세 스케일부터 조잡한 스케일까지 감쇠시키면서 스펙트럼 일치 가우스 노이즈를 주입하는 과정)을 설계했습니다. 여기서 스케일은 확산 역학의 명시적인 좌표가 됩니다. 동일하게 학습된 역방향 프로세스는 시작 시간 단계를 변경하여 이미지 생성과 연속적인 초해상도를 수행하며, $ extit{작업별 특화 아키텍처, 조건부 분기, 분류기 없는 가이드, 스케일 계수에 따른 재학습이 필요하지 않습니다}$. 실험 결과, SKILD는 unconditional CIFAR-10 데이터셋에서 FID 점수 2.65 및 Inception Score 9.63을 달성했으며, 단일 unconditional 체크포인트를 사용하여 ImageNet 데이터셋에 대해 2배 ~ 8배의 초해상도를 수행하여 다양한 시각적 지표에서 조건부 모델보다 우수한 성능을 보였습니다. 또한, SKILD는 연결된 네 점 상관관계가 실제 값과 밀접하게 일치하는 Ising 모델을 성공적으로 복원했습니다.

Original Abstract

Creating images from noise is image generation; reconstructing fine details from coarse inputs is super-resolution. Despite their practical differences, both can be understood as reversing information loss across scales. We introduce $\textbf{SKILD}$, a $\textbf{S}$cale-invariant $\textbf{K}$-Space $\textbf{I}$mage $\textbf{L}$earning $\textbf{D}$iffusion model that unifies generation and continuous super-resolution within a single unconditional framework. Both natural images and critical physical systems exhibit scale invariance, and we leverage it to design a forward process that attenuates image content from fine to coarse scales while injecting spectrum-matched Gaussian noise, making scale an explicit coordinate of the diffusion dynamics. The same trained reverse process performs generation and continuous super-resolution by varying only the starting timestep: $\textit{no task-specific architecture, no conditioning branch, no classifier-free guidance, no retraining per scale factor}$. Empirically, SKILD reaches FID $2.65$ and Inception Score $9.63$ on unconditional CIFAR-10, performs $2\times$--$8\times$ super-resolution on ImageNet from a single unconditional checkpoint while outperforming conditional models across perceptual metrics, and reconstructs critical Ising models whose connected four-point correlations closely track the ground truth.

0 Citations
0 Influential
7.5 Altmetric
37.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!