샘플 기반 변분 추론을 활용한 소규모 단계 생성 모델 정렬
Aligning Few-Step Generative Models by Amortizing Sample-based Variational Inference
몇 단계로 구성된 생성 모델을 정렬하는 것은 어려운 과제이며, 기존의 정렬 프레임워크는 일반적으로 다음과 같은 제한적인 가정을 필요로 합니다. 즉, 계산 가능한 likelihood, 특정 ODE/SDE 솔버 또는 특정 모델 패밀리입니다. 본 논문에서는 FAV (Few-step Alignment via Variational Inference)라는 일반적인 정렬 프레임워크를 소개합니다. FAV은 생성기와 참조 분포에 대한 샘플 데이터만 필요로 합니다. 우리는 정렬을 보상 가중치를 적용한 분포에서 샘플링하는 문제로 정의하며, 이 분포는 참조 분포를 기준으로 합니다. 또한, Stein Variational Gradient Descent를 활용하여 샘플 기반 변분 추론을 수행하고, 파티클 업데이트를 고정점 회귀를 통해 생성 모델의 파라미터에 반영합니다. FAV은 로봇 조작 및 이미지 생성 정렬이라는 두 가지 분야에서 평가되었습니다. 로봇 조작을 위한 정책 정렬 작업에서는 FAV이 기존의 정책 추출 기반 방법보다 56개의 오프라인 작업과 30개의 오프라인-온라인 강화 학습 작업에서 더 우수한 성능을 보였습니다. 이미지 생성기 정렬 작업에서는 GAN, drifting 모델, consistency 모델 및 flow map과 같은 다양한 소규모 단계 기반 모델을 fine-tuning했으며, ImageNet-$256$에서 1024$^2$ 크기의 텍스트-이미지 합성까지 확장했습니다. 관련 코드는 https://github.com/Jaewoopudding/FAV 에서 확인할 수 있습니다.
Aligning a few-step generative model is challenging, since existing alignment frameworks typically rely on restrictive assumptions: a tractable likelihood, a specific ODE/SDE solver, or a particular model family. We introduce FAV, Few-step Generative Models Alignment via Sample-based Variational Inference, a general alignment framework that requires only sample access to the generator and the reference distribution. We cast alignment as sampling from a reward-tilted distribution anchored to a reference distribution. We leverage Stein Variational Gradient Descent as a sample-based variational inference scheme and amortize its particle updates into the generator parameters via fixed-point regression. We evaluate FAV on two domains: robotics manipulation and image generator alignment. On generative policy alignment for robotic manipulation, FAV outperforms prevailing policy extraction baselines across 56 offline and 30 offline-to-online RL tasks. For image generator alignment, FAV fine-tunes diverse few-step backbones, including GAN, drifting model, consistency models, and flow maps, scaling from ImageNet-$256$ to 1024$^2$ text-to-image synthesis. Code is available at https://github.com/Jaewoopudding/FAV.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.