에너지 기반 플로우 매칭
Energy-Guided Flow Matching
픽셀 공간 생성 모델은 손실 압축을 피할 수 있지만, 고차원 공간에서 전역 구조와 미세한 디테일을 동시에 학습해야 합니다. 일반적인 플로우 매칭 방법은 노이즈를 고정된 깨끗 이미지 목표 지점으로 보간하지만, 스펙트럼 변화는 암묵적으로 학습됩니다. 본 논문에서는 에너지 기반 플로우 매칭(EG-FM)을 제안합니다. EG-FM은 이동하는 목표 지점을 사용하여 거시적에서 미세적인 생성 경로를 명시적으로 모델링합니다. 구체적으로, EG-FM은 고정된 목표 지점을 저주파 이미지에서 깨끗한 이미지로 부드럽게 변화하는 히트 커널 필터링된 목표 지점으로 대체합니다. 이미지별 에너지 기반 스케줄링을 통해 이동하는 목표 지점에 포함된 고주파 신호의 비율을 조절하여 플로우 매칭에서의 속도 벡터를 재조정합니다. 우리의 프레임워크는 백본 네트워크나 학습 데이터에 대한 어떠한 추가적인 조정도 필요하지 않으며, 학습 및 추론 단계에서 미미한 비용만 발생합니다. 실험 결과, EG-FM은 ImageNet 클래스 조건 이미지 생성 작업에서 $256 imes 256$ 해상도에서 더 적은 에포크로 일관되게 낮은 FID 값을 달성했습니다. 200 에포크에서 FID 값이 1.55이고, 600 에포크에서 1.45를 기록했습니다. 또한, $512 imes 512$ 해상도로 생성 작업을 추가 학습한 결과, 단 40 에포크의 고해상도 적응을 통해 FID 값이 1.58로 향상되었습니다. 더불어, EG-FM을 텍스트-이미지 생성 작업에 적용하여 GenEval 점수 0.85 및 DPG-Bench 점수 83.9를 달성했습니다. 코드: https://github.com/ysng123/EG-FM
Pixel-space generative models bypass lossy latent compression, yet necessitate joint learning of global structure and fine-grained details in a high-dimensional space. Standard flow matching interpolates noise toward a fixed clean-image endpoint, leaving the spectral evolution to be learned implicitly. In this paper, we introduce Energy-Guided Flow Matching(EG-FM) that explicitly models a coarse-to-fine generative trajectory by moving endpoint. Specifically, EG-FM replaces the fixed endpoint with a heat-kernel-filtered endpoint that evolves smoothly from low-frequency image to clean image. The fraction of high-frequency signal in moving endpoint is released by an image-specific energy-guided scheduling, leading to the re-targeting of velocity in flow matching. Our framework requires no adaptation of the backbone and training data, bringing negligible cost on the training and inference stages. In our experiment, EG-FM consistently achieves lower FID on the ImageNet class-conditional image generation task at $256 \times 256$ with fewer epochs, reaching an FID of 1.55 at 200 epochs and 1.45 at 600 epochs. We continue training the generation task on the setting of $512 \times 512$ resolution, yielding a FID of 1.58 after only 40 high-resolution adaptation epochs. Furthermore, we transfer EG-FM on text-to-image generation and achieve 0.85 on GenEval score and 83.9 on DPG-Bench. Code is available at https://github.com/ysng123/EG-FM.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.