TARS: 타임스텝 기반 데이터 스케일링을 통한 3차원 정보 불필요 비디오 재촬영
TARS: Timestep-Aware Data Scaling for 3D-Free Video Re-Shooting
비디오 재촬영은 제어 가능한 카메라 움직임과 시점을 사용하여 영상을 재생성하는 것을 목표로 합니다. 기존 방법들은 명시적인 3차원 정보를 필요로 하는데, 이는 복원 품질에 의해 제한되며 종종 새로운 영역을 합성할 때 성능이 저하됩니다. 또한, 서로 다른 카메라 경로를 가진 쌍으로 연결된 비디오 데이터를 사용하는 방법은 데이터의 희소성 때문에 일반화에 어려움을 겪습니다. 본 연구에서는 텍스트 기반의 의미론적 시점 제어를 통해 비디오 재촬영 문제를 다시 접근하며, 촬영 규모, 시야각 및 1인칭/3인칭 관점을 제어할 수 있도록 합니다. 이를 위해 3차원 정보 없이 작동하는 비디오 재촬영 패러다임인 TARS를 제안합니다. 타임스텝별 민감도 분석 결과, 카메라 움직임은 주로 노이즈가 큰 단계에서 결정되며, 이때 거친 시공간 구조가 형성됩니다. 이러한 통찰력을 바탕으로, TARS는 쌍으로 연결된 재촬영 데이터나 3차원 복원 없이 카메라 동역학과 기본적인 시각적 표현을 학습하는 자기 지도 학습 방식을 도입합니다. 데이터 스케일링 및 텍스트-카메라 결합 조건을 통해 TARS는 강력한 카메라 및 시점 제어를 지원하며, 큰 카메라 움직임 하에서 원본 시야 범위를 벗어난 영역을 자연스럽게 합성할 수 있습니다. 또한, 역각 재촬영과 관점 전환도 가능합니다. 광범위한 실험 결과, TARS가 기존 방법보다 더 정확하고 일관된 카메라 제어를 제공하는 것으로 나타났습니다. 프로젝트 페이지: https://ymlinfeng.github.io/TARS.github.io/
Video re-shooting aims to regenerate videos with controllable camera motion and viewpoint. Existing methods rely on explicit 3D priors, which are limited by reconstruction quality and often perform poorly when synthesizing previously unseen regions, or on paired videos with different camera trajectories, whose scarcity hinders generalization. We revisit video re-shooting through text-driven semantic viewpoint specification, enabling control over shot scale, viewing angle, and first-/third-person perspective. To this end, we propose TARS, a 3D-free video re-shooting paradigm. Timestep-wise sensitivity analysis reveals that camera motion is primarily established during high-noise stages, where coarse spatiotemporal structures are formed. Based on this insight, we introduce self-supervised training to learn camera dynamics and fundamental visual representations without paired re-shooting data or 3D reconstruction. Through data scaling and joint textual-camera conditioning, TARS supports robust camera and viewpoint control, plausibly synthesizing regions beyond the source view under large camera motions while enabling reverse-angle re-shooting and perspective switching. Extensive experiments show that TARS provides more accurate and temporally consistent camera control than prior methods. Project Page: https://ymlinfeng.github.io/TARS.github.io/
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.