2607.26811v1 Jul 29, 2026 cs.CV

DistillAlign: 자기회귀 비디오 증류 과정에서 모드 커버리지와 모드 탐색의 조화

DistillAlign: Coordinating Mode Covering and Mode Seeking in Autoregressive Video Distillation

Jiaxing Li
Jiaxing Li
Citations: 409
h-index: 3
Kai Zou
Kai Zou
Citations: 234
h-index: 3
Cindy Zhou
Cindy Zhou
Citations: 0
h-index: 0
Kaichen Huang
Kaichen Huang
Citations: 22
h-index: 2
Junyao Gao
Junyao Gao
Citations: 348
h-index: 7
Zile Wang
Zile Wang
Citations: 57
h-index: 5
Yang Liu
Yang Liu
Citations: 0
h-index: 0
Bin Liu
Bin Liu
Citations: 313
h-index: 7
Bo An
Bo An
Citations: 45
h-index: 2
Yangguang Li
Yangguang Li
Citations: 316
h-index: 8

기존의 자기회귀 비디오 증류 방법은 일반적으로 분포 매칭 증류(DMD) 기반의 다단계 파이프라인을 사용합니다. 그러나 이러한 방법들은 초기화 단계와 DMD 단계를 분리하여, 서로 다른 목표 분포를 추구하며, 중간 단계 학생 모델의 성능을 주로 시각적 점수(예: VBench)로 평가합니다. 본 논문에서는 이러한 설계를 분포론적인 관점에서 재검토합니다. 분포 매칭 손실은 모드 탐색에 중점을 두므로, 좋은 초기화는 대상 DMD 선생님 모델의 모드 커버리지를 일치시켜야 하며, 단순히 높은 품질을 추구해서는 안 됩니다. 이를 분석하기 위해, 학생 모델과 선생님 모델의 분포 간의 정밀도와 커버리지를 측정하는 분포 평가 프로토콜을 도입했습니다. 이는 시각적 점수로 숨겨진 차이점을 드러냅니다. 일부 초기화는 높은 정밀도를 달성하지만 낮은 커버리지를 가지며, 이는 최적의 개선으로 이어지지 않습니다. 반면, 모드 커버리지가 높은 초기화는 더 넓은 범위를 유지합니다. 또한, 대상 분포가 일치하더라도 DMD의 역 KL 목적 함수는 훈련 후반 단계에서 학생 모델을 선생님 모델의 고 확률 영역으로 유도하여 커버리지와 다양성을 감소시킬 수 있습니다. 이러한 문제를 해결하기 위해, 우리는 DMD의 모드 탐색 목표와 Consistency Distillation 기반의 모드 커버리지 제약을 결합한 통합 증류 방법을 제안합니다. 실험 결과는 우리의 방법이 생성 품질, 커버리지 및 다양성을 향상시키며, 특히 Wan-1.3B DMD 선생님 모델을 사용하는 경우에도 Wan-14B로 개선된 기준 모델보다 우수한 성능을 보인다는 것을 보여줍니다. 이는 자기회귀 비디오 증류에서 분포 정렬의 중요성을 강조합니다.

Original Abstract

Existing autoregressive video distillation methods commonly adopt a Distribution Matching Distillation (DMD)-based multi-stage pipeline. However, they typically decouple the initialization and DMD stages -- which then pursue different target distributions -- and judge the intermediate student mainly by visual scores such as VBench. In this paper, we revisit this design from a distributional perspective. Given the mode-seeking nature of the distribution matching loss, a good initialization should match the mode coverage of the target DMD teacher, rather than merely pursuing high quality. To analyze this, we introduce a distributional evaluation protocol that measures precision and coverage between student and teacher distributions in a shared latent space. It exposes differences hidden by visual scores: some initializations reach high precision but low coverage, leading to suboptimal refinement, while mode-covering ones preserve broader support. Furthermore, even when the target distributions are aligned, DMD's reverse-KL objective can still drive the student toward high-probability teacher regions in late training, reducing coverage and diversity. To address this, we propose joint distillation, which combines DMD's mode-seeking objective with a Consistency Distillation-based mode-covering constraint. Experiments show that our method improves generation quality, coverage, and diversity; notably, even with a Wan-1.3B DMD teacher, it outperforms baselines refined with Wan-14B, underscoring the importance of distributional alignment in autoregressive video distillation.

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!