2607.26004v1 Jul 28, 2026 cs.CV

병렬 디코딩 증류를 통한 고속 이미지 및 비디오 생성

Parallel Decoding Distillation for Fast Image and Video Generation

Julius Berner
Julius Berner
Citations: 81
h-index: 4
Chao Liu
Chao Liu
Citations: 269
h-index: 6
Arash Vahdat
Arash Vahdat
Citations: 7,073
h-index: 30
Neta Shaul
Neta Shaul
Citations: 889
h-index: 9

비디오 확산 또는 흐름 모델에서의 생성이 계산 비용이 많이 드는 이유는 느리고 반복적인 샘플링 과정 때문입니다. 현재 최고 성능을 보이는 가속화 방법은 주로 변분 스코어 증류(VSD)와 적대적 손실을 사용하여 확산 모델을 소수의 단계로 구성된 생성기로 변환하는 데 의존합니다. 이러한 방법들은 고품질 비디오 생성을 달성하지만, 훈련 과정에서 최적화하기 매우 어렵고 모드 콜랩스 문제를 일으켜 비디오의 다양성을 저해하고 움직임 표현에 부족함을 초래합니다. 본 논문에서는 병렬 디코딩 증류(PDD)를 소개하며, 이는 확산 및 흐름 매칭 모델의 빠른 추론을 위한 단순화되고 확장 가능한 경로 기반 증류 방법입니다. 제안하는 아키텍처 및 훈련 절차는 모든 사전 훈련된 모델과 호환되며, 다양한 함수 평가 횟수(NFE)로 샘플링을 지원합니다. PDD는 네트워크 평가 한 번당 여러 개의 디노이징 단계를 예측하여 생성을 가속화합니다. 개념적으로, JVPs 또는 유한 차분 근사를 사용하지 않고 평균 속도의 표현을 학습합니다. 제안하는 방법은 LTX-2.3 텍스트-비디오/오디오, Wan 14B 텍스트-비디오, Qwen-Image 텍스트-이미지 모델에서 각각 4~8개의 NFE를 사용하여 최고 성능을 달성했습니다. 또한, PDD는 생성된 비디오의 다양성을 크게 향상시킵니다.

Original Abstract

Generation in video diffusion or flow models is computationally expensive due to the slow and iterative sampling process. Current state-of-the-art (SOTA) acceleration methods heavily rely on variational score distillation (VSD) and adversarial losses to distill diffusion models into few-step generators. Albeit achieving high-quality video generation, these training losses are notoriously hard to optimize and suffer from mode collapse, leading to loss of video diversity and lack of motion. In this paper, we introduce Parallel Decoding Distillation (PDD), a simplified and scalable trajectory-based distillation method for fast inference of diffusion and flow matching models. Our architecture and training procedure are compatible with any pre-trained model and support sampling with a varying number of function evaluations (NFE). PDD accelerates generation by predicting multiple denoising steps per network evaluation. Conceptually, it learns a representation of the mean velocity without regressing its derivative using JVPs or finite-difference approximations. Our method achieves SOTA performance with 4-8 NFE on LTX-2.3 Text-to-Video/Audio, Wan 14B Text-to-Video, and Qwen-Image Text-to-Image. Moreover, PDD presents a significant improvement in generated video diversity.

0 Citations
0 Influential
15 Altmetric
75.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!