2608.13556v1 Aug 13, 2026 cs.CV

V-RAE: 비디오 잠재 공간 재구상 - 생성 모델을 위한 접근 방식

V-RAE: Rethinking Video Latent Spaces for Generation

Hao Fei
Hao Fei
Citations: 594
h-index: 8
Minghui Guo
Minghui Guo
Citations: 235
h-index: 8
Shengqiong Wu
Shengqiong Wu
Citations: 3,334
h-index: 28

비디오 생성은 일반적으로 자동 인코더를 사용하여 생성 모델이 작동하는 압축된 공간을 정의합니다. 비디오 자동 인코더 아키텍처는 크게 발전했지만, 여전히 픽셀 수준의 복원에 최적화되어 있으며 제한적인 고수준 의미론적 구조를 제공합니다. 그러나 복원 성능에 최적화된 잠재 공간이 생성 모델링에 적합할 필요는 없습니다. 본 논문에서는 V-RAE라는 비디오 표현 자동 인코더를 제안합니다. V-RAE는 정제된(frozen) 시각 기반 모델의 표현 위에 압축된 생성 잠재 공간을 구축합니다. 경량화된 시간 축 풀링 모듈은 시간적 중복성을 제거하면서 의미론적 구조를 유지하며, 비디오 디코더는 압축된 특징으로부터 연속적인 움직임을 재구성합니다. V-RAE는 네 가지 대표적인 정제된 인코더를 사용하여 비디오 복원, 의미론적 분석 및 조건부 생성에 대해 평가했습니다. V-RAE는 K600에서 2.13의 rFVD 값을 달성하여, 평가된 모든 대규모 사전 학습된 비디오 VAE보다 우수한 성능을 보였습니다. 또한 V-RAE의 잠재 표현은 기존의 비디오 토크나이저 잠재 표현보다 훨씬 더 많은 의미론적 정보를 유지합니다. 동일한 생성 설정을 사용했을 때, 최상의 변형 모델은 UCF101 및 K600에서 각각 117.86과 19.16의 gFVD 점수를 달성했으며, 수렴 속도가 최대 6배 더 빠릅니다. 또한 복원 품질만으로는 생성 유용성을 평가하기에 충분하지 않으며, 다운스트림 생성 품질과 더욱 신뢰할 수 있는 상관관계를 보이는 시간적 일관성 진단 도구인 tFVD를 제안합니다. 비디오 생성을 넘어, V-RAE는 동일한 예측 설정 하에서 Cityscapes 데이터셋에서 Wan 2.2 VAE 잠재 공간보다 미래 비디오 예측 성능을 향상시킵니다. 종합적으로 볼 때, 본 연구 결과는 정제된 의미론적 표현이 비디오 복원, 생성 및 예측 모델링을 지원할 수 있음을 보여줍니다. 프로젝트 페이지: https://v-rae.github.io/.

Original Abstract

Latent video generation relies on autoencoders to define a compact space in which generative models operate. Although video autoencoder architectures have evolved substantially, their latent spaces are still optimized primarily for pixel-level reconstruction and provide limited high-level semantic organization. A reconstruction-optimal latent space, however, need not be well suited to generative modeling. We propose V-RAE, a video representation autoencoder that builds compact generative latents on top of frozen vision foundation model representations. A lightweight temporal pooling module removes temporal redundancy while preserving semantic structure, and a video decoder reconstructs continuous motion from the compressed features. We evaluate V-RAE with four representative frozen encoders on video reconstruction, semantic probing, and class-conditional generation. V-RAE achieves 2.13 rFVD on K600, outperforming all evaluated large-scale pretrained video VAEs. Its latents retain substantially more semantic information than conventional video tokenizer latents. Under matched generation settings, our best variant achieves gFVD scores of 117.86 and 19.16 on UCF101 and K600, respectively, while converging up to 6x faster}. We further show that reconstruction quality alone is insufficient to characterize generative utility and introduce tFVD, a temporal-coherence diagnostic that correlates more reliably with downstream generation quality. Beyond video generation, V-RAE also improves future video prediction on Cityscapes over the Wan 2.2 VAE latent space under matched prediction settings. Taken together, the experiments show that frozen semantic representations can support video reconstruction, generation, and predictive modeling. The project page: https://v-rae.github.io/.

0 Citations
0 Influential
14 Altmetric
70.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!