2607.26860v1 Jul 29, 2026 cs.LG

시각적 생성에 대한 암모트화된 모멘트 매칭

Amortized Moment Matching for Visual Generation

Xintao Wang
Xintao Wang
Citations: 2,117
h-index: 21
Wenze Liu
Wenze Liu
Citations: 235
h-index: 4
Pengfei Wan
Pengfei Wan
Citations: 2,981
h-index: 27
Xiangyu Yue
Xiangyu Yue
Citations: 105
h-index: 4

본 논문에서는 신경망을 사용하여 데이터의 통계적 특징(모멘트)을 학습하고, 이를 분포 기반의 학습 신호로 활용하는 암모트화된 모멘트 매칭 방법을 제안합니다. 확산 모델의 노이즈 제거 과정을 다항식 투영으로 표현함으로써, 일반적인 암모트화 프레임워크를 구축하며, n차 투영은 최대 (n+1) 차까지의 데이터 모멘트를 명시적으로 식별한다는 것을 밝힙니다. 이로부터, 계산 가능한 선형(affine) 모델을 기반으로 Amortized Fréchet Distance (AMFD) 손실 함수를 정의합니다. 기존 FD-loss는 명시적인 주변 분포 모멘트 계산에 의존하는 반면, AMFD는 대규모 데이터셋에 효율적으로 적용될 수 있는 반복적이고 행렬 연산을 사용하지 않는 최적화 과정을 통해 조건부 모멘트를 동적으로 학습합니다. 전역 표현 특징을 사용할 때, AMFD는 강력한 후처리 목표 함수 역할을 하며, 실제 실험 결과에서 신경망 기반의 AMFD는 정확한 통계적 매칭보다 더 안정적인 학습 과정을 제공하며, FDr⁶ 지표에서 기존 FD 모델을 크게 능가하고 ImageNet 데이터셋에 대한 단일 단계 생성 성능이 우수합니다. 또한, AMFD는 원본 생성 공간 내에서의 탐색을 가능하게 하며, 처음 두 개의 모멘트가 강한 의미론적 구조를 가진 공간에서만 목표 분포를 식별할 수 있음을 시사합니다. 마지막으로, 텍스트-이미지 생성 모델에 적용했을 때, AMFD의 조건 인식 특성은 지시 사항 준수 능력을 크게 향상시켜, 단일 단계 모델이 FLUX.2 [klein] 4B 모델과 같은 다단계 모델을 GenEval 벤치마크에서 능가하는 성능을 보여주면서도 PickScore에서는 동등한 성능을 달성합니다. 코드 및 체크포인트는 https://github.com/poppuppy/amfd 에서 제공됩니다.

Original Abstract

We propose amortized moment matching, utilizing neural networks to learn data moments as distributional training signals. By casting diffusion denoisers through polynomial projections, we establish a general framework for moment amortization, revealing that an $n$-th degree projection explicitly identifies data moments up to order $n+1$. Derived from the tractable affine case, we instantiate the Amortized Fréchet Distance (AMFD) loss. Unlike FD-loss which relies on explicit marginal moment calculations, AMFD is able to dynamically learn conditional moments via an alternating, matrix-free optimization pipeline that effortlessly scales to high-dimensional data. When operating on global representation features, AMFD serves as a powerful post-training objective; empirically, its neural formulation yields more robust training dynamics than exact statistical matching, substantially surpassing the FD baseline on the FDr$^6$ metric and achieving superior one-step generation on ImageNet. Furthermore, it unlocks direct exploration within native generative spaces, suggesting that the first two moments can identify target distributions only in spaces with strong semantics. Finally, when scaled to text-to-image generation, the condition-aware nature of AMFD unlocks massive gains in instruction-following capabilities, enabling our one-step models to outperform their multi-step FLUX.2 [klein] 4B teachers on the GenEval benchmark while achieving on-par performance on PickScore. Code and checkpoints are available at https://github.com/poppuppy/amfd.

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!