2603.21864v1 Mar 23, 2026 cs.CV

적응형 비디오 증류: 소량 단계 생성 시 과포화 및 시간적 붕괴 완화

Adaptive Video Distillation: Mitigating Oversaturation and Temporal Collapse in Few-Step Generation

Yuyang You
Yuyang You
Citations: 30
h-index: 2
Yongzhi Li
Yongzhi Li
Citations: 120
h-index: 4
Jiahui Li
Jiahui Li
Citations: 34
h-index: 2
Yadong Mu
Yadong Mu
Citations: 385
h-index: 7
Quan Chen
Quan Chen
Citations: 89
h-index: 4
Peng Jiang
Peng Jiang
Citations: 73
h-index: 4

최근 비디오 생성은 생성형 AI 분야의 핵심 과제로 부상했습니다. 그러나 비디오 합성의 상당한 계산 비용은 모델 증류를 효율적인 배포를 위한 중요한 기술로 만듭니다. 그 중요성에도 불구하고, 비디오 확산 모델을 위해 특별히 설계된 방법은 부족합니다. 기존 접근 방식은 종종 이미지 증류 기술을 직접적으로 적용하며, 이는 과포화, 시간적 불일치 및 모드 붕괴와 같은 문제점을 야기합니다. 이러한 문제점을 해결하기 위해, 비디오 확산 모델을 위해 특별히 설계된 새로운 증류 프레임워크를 제안합니다. 이 프레임워크의 핵심 혁신은 다음과 같습니다: (1) 과도한 분포 변화로 인해 발생하는 문제점을 방지하기 위해 공간적 지도 가중치를 동적으로 조정하는 적응형 회귀 손실; (2) 시간적 붕괴를 방지하고, 부드럽고 물리적으로 타당한 샘플링 경로를 촉진하는 시간적 정규화 손실; (3) 샘플링 오버헤드를 줄이면서도 인지적 품질을 유지하는 추론 시간 프레임 보간 전략. VBench 및 VBench2 벤치마크에 대한 광범위한 실험 및 분석 연구는 제안된 방법이 안정적인 소량 단계 비디오 생성을 달성하며, 인지적 충실도와 움직임 사실감을 크게 향상시킨다는 것을 보여줍니다. 또한, 여러 지표에서 기존 증류 기준보다 일관되게 우수한 성능을 보입니다.

Original Abstract

Video generation has recently emerged as a central task in the field of generative AI. However, the substantial computational cost inherent in video synthesis makes model distillation a critical technique for efficient deployment. Despite its significance, there is a scarcity of methods specifically designed for video diffusion models. Prevailing approaches often directly adapt image distillation techniques, which frequently lead to artifacts such as oversaturation, temporal inconsistency, and mode collapse. To address these challenges, we propose a novel distillation framework tailored specifically for video diffusion models. Its core innovations include: (1) an adaptive regression loss that dynamically adjusts spatial supervision weights to prevent artifacts arising from excessive distribution shifts; (2) a temporal regularization loss to counteract temporal collapse, promoting smooth and physically plausible sampling trajectories; and (3) an inference-time frame interpolation strategy that reduces sampling overhead while preserving perceptual quality. Extensive experiments and ablation studies on the VBench and VBench2 benchmarks demonstrate that our method achieves stable few-step video synthesis, significantly enhancing perceptual fidelity and motion realism. It consistently outperforms existing distillation baselines across multiple metrics.

3 Citations
0 Influential
3.5 Altmetric
20.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!