2608.04818v1 Aug 05, 2026 cs.CV

Interval Denoiser를 활용한 픽셀 평균 흐름 재고찰

Rethinking Pixel Mean Flows via Interval Denoiser

Aibek Alanov
Aibek Alanov
Citations: 329
h-index: 10
A.M. Zaytsev
A.M. Zaytsev
Citations: 0
h-index: 0
Dmitry Baranchuk
Dmitry Baranchuk
Citations: 52
h-index: 3
Alexander Korotin
Alexander Korotin
Citations: 1,175
h-index: 18

최근의 확산 모델 및 흐름 기반 모델은 다단계 샘플링으로 인한 계산 부담과 외부 오토인코더의 복구 병목 현상을 해결하기 위해 점진적으로 단계 수를 줄이고 잠재 공간을 사용하지 않는 생성 방식으로 발전하고 있습니다. 본 논문에서는 잠재 공간 없이 작동하는 엄밀한 이론적 프레임워크인 Interval Denoiser를 제안합니다. Interval Denoiser는 흐름 매칭 ODE로부터 직접 파생되었으며, 중간 경로 상태에 대한 정확한 분석적 맵핑을 제공합니다. 기존 방식과 달리, 본 논문에서 제시된 예측 결과는 임의의 시간 구간에서 저차원 매니폴드 상에 존재하며, 이를 통해 네트워크가 픽셀 데이터를 직접 사용하여 회귀 문제를 해결할 수 있습니다. 또한, 경험적인 대수 변환을 피함으로써 순수한 시간 미분을 정확하게 분리하여 편향된 기울기 계산을 방지하고 정확한 1차 최적화를 보장합니다. 이러한 객체 함수 분석을 바탕으로, 본 논문에서는 잔여 클리핑 및 시간 샘플링 커리큘럼을 적용하여 효율적인 장기간 학습을 가능하게 하고, 적은 단계 수에서의 성능을 향상시켰습니다. ImageNet 256x256 데이터셋에서 처음부터 학습한 모델은 감성적 손실(perceptual loss) 없이 단일 단계(1-NFE)에서 FID 점수가 4.55이고, 두 단계(2-NFE)에서 3.98을 달성했습니다.

Original Abstract

Modern diffusion and flow-based models are increasingly moving toward few-step, latent-free generation to bypass the computational overhead of multi-step sampling and the reconstruction bottlenecks of external autoencoders. We propose the Interval Denoiser, a theoretically rigorous framework for latent-free generation. Derived directly from the flow matching ODE, it establishes an exact analytical mapping for intermediate trajectory states. Unlike prior formulations, our prediction is shown to reside on a low-dimensional manifold across any time interval, making the regression tractable for a network operating directly on pixels. Furthermore, by avoiding empirical algebraic substitutions, our formulation correctly isolates the pure time derivative to prevent biased gradient evaluations and ensure exact first-order optimization. By analyzing this objective, we equip our framework with residual clipping and a time-sampling curriculum, enabling effective long-interval training and improving few-step performance. Trained from scratch on ImageNet 256x256, our model achieves an FID of 4.55 in one step (1-NFE) and 3.98 in two steps (2-NFE) without perceptual losses.

0 Citations
0 Influential
9 Altmetric
45.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!