2606.31050v1 Jun 30, 2026 cs.CV

예측 가능 미분 렌더링을 이용한 비디오 동역학 학습

Learning Video Dynamics with Predictive Differentiable Rendering

Cheng Tan
Cheng Tan
Citations: 232
h-index: 9
Tian Zhou
Tian Zhou
Citations: 153
h-index: 8
Liang Sun
Liang Sun
Citations: 94
h-index: 6
Rong Jin
Rong Jin
Citations: 4,907
h-index: 12
SouYoung Jin
SouYoung Jin
Citations: 37
h-index: 2
Yifan Hu
Yifan Hu
Citations: 11
h-index: 1
Yujin Tang
Yujin Tang
Citations: 13
h-index: 3
Xin Lin
Xin Lin
Citations: 0
h-index: 0

고품질의 미래 장면을 정확하게 예측하는 방법은 무엇일까요? 시각적 세계는 본질적으로 연속적이̆지만, 기존의 결정론적인 비디오 예측 모델들은 이산적인 픽셀 공간에서 작동하며, 주로 픽셀 단위의 평균 제곱 오차(MSE)를 통해 최적화됩니다. 이는 종종 과도하게 부드러운 예측 결과를 초래하고, 미세한 시각적 디테일이 부족해지는 문제를 야기합니다. 이러한 한계점을 해결하기 위해, 본 연구에서는 이산적인 표현과 연속적인 표현 사이의 간극을 좁히는 새로운 엔드-투-엔드 비디오 예측 패러다임인 Predictive Differentiable Rendering (PDR)을 제안합니다. 최근 3D 가우시안 스플래팅 기반 3D 재구성 기술의 발전에 영감을 받아, 우리는 2D 가우시안 표현을 기반으로 하는 경량화되고 쉽게 통합 가능한 어댑터인 PredGS를 소개합니다. PredGS는 기존의 픽셀 공간 예측 모델과 원활하게 통합되어 계산 비용은 거의 증가하지 않으면서도 공간적 디테일 보존 능력을 크게 향상시킵니다. 또한, 우리는 CUDA 가속을 지원하는 미분 가능한 2D 가우시안 렌더러인 predgsplat를 개발했습니다. 각 가우시안은 5 + C개의 학습 가능한 파라미터(위치, 크기, 회전 및 C 채널의 진폭)로 정의되며, 기준 모델보다 최대 10배 빠른 렌더링 속도를 제공합니다. L1 손실과 SSIM 손실을 결합하여 최적화된 PDR은 MSE 손실이 가진 고유한 흐릿함 문제를 극복하고 예측 성능을 크게 향상시킵니다. TaxiBJ, WeatherBench, KTH 및 Human3.6M 등 다양한 실제 데이터셋에 대한 광범위한 실험 결과는 PDR이 기존 방법들을 능가하며, 뛰어난 디테일 보존 능력, 시각적 충실도 및 예측 정확도를 제공한다는 것을 보여줍니다.

Original Abstract

How to accurately predict a high-fidelity future world? While the visual world is inherently continuous, existing deterministic video prediction models operate in discrete pixel space and are mainly optimized with pixel-wise mean squared error (MSE), which often leads to over-smoothed predictions and a lack of fine-grained visual details. To address these limitations, we propose Predictive Differentiable Rendering (PDR), a novel end-to-end video prediction paradigm that bridges the gap between discrete and continuous representations. Inspired by recent progress in 3D reconstruction with 3D Gaussian Splatting, we introduce PredGS, a lightweight and plug-and-play adapter based on 2D Gaussian representation, which could be seamlessly integrated with existing pixel space predictors, significantly improving spatial detail preservation with negligible computational overhead. Furthermore, we develop predgsplat, a CUDA-accelerated differentiable 2D Gaussian renderer supporting arbitrary channels. Each Gaussian is defined by 5 + C learnable parameters (position, scale, rotation, and C channel amplitudes) and achieves up to 10x faster rendering than the baseline. Optimized by a combined L1 and SSIM loss, PDR overcomes the inherent blurring tendencies of MSE Loss, significantly enhancing the prediction performance. Extensive experiments on diverse real-world benchmarks, including TaxiBJ, WeatherBench, KTH, and Human3.6M, demonstrate that PDR consistently surpasses existing methods, delivering superior detail preservation, visual fidelity, and predictive accuracy.

0 Citations
0 Influential
6 Altmetric
30.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!