JoyAI-Video-Edit: 자기 회귀 확산 모델을 활용한 실시간, 개방형 비디오 편집
JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion
실시간 비디오 편집은 낮은 지연 시간으로 원본의 충실도를 유지하고 장기적인 시간적 일관성을 확보하면서 제한된 계산 자원을 사용해야 합니다. 본 논문에서는 미래 프레임이나 미리 정의된 비디오 길이에 의존하지 않고 실시간, 개방형 비디오 편집을 가능하게 하는 160억 개의 파라미터를 가진 자기 회귀 확산 모델 기반 프레임워크인 JoyAI-Video-Edit를 제안합니다. 저희 방법은 청크별 자기 회귀 적응, 소스 기반 분포 일치 증류(SA-DMD), 그리고 장기 자기 회귀 증류를 결합하여 학습과 추론 간의 불일치를 줄이고, 두 단계 생성 과정에서 원본의 충실도를 유지하며, 누적된 시간적 드리프트를 완화합니다. 광범위한 자동 및 인간 평가 결과, JoyAI-Video-Edit는 기존 스트리밍 편집 도구를 크게 능가하며, 짧은 비디오와 긴 비디오 모두에서 강력한 오프라인 시스템과 경쟁력 있는 성능을 보여줍니다. 전체 시스템은 단일 Nvidia B200 GPU에서 약 30 FPS의 속도로 720p 비디오 편집을 수행할 수 있습니다. 코드는 https://github.com/jd-opensource/JoyAI-Video-Edit 에서 확인할 수 있습니다.
Real-time video editing requires low-latency causal generation with bounded computational resources while preserving source fidelity and long-term temporal consistency. We present JoyAI-Video-Edit, a 16B-parameter autoregressive diffusion framework for real-time, open-ended video editing without access to future frames or a predefined video duration. Our method combines chunk-wise autoregressive adaptation, Source-Anchored Distribution Matching Distillation (SA-DMD), and Long-Horizon Autoregressive Distillation to reduce train--inference mismatch, preserve source fidelity during two-step generation, and mitigate accumulated temporal drift. Extensive automatic and human evaluations show that JoyAI-Video-Edit substantially outperforms existing streaming editors and remains competitive with strong offline systems on both short and long videos. The complete system achieves end-to-end 720p video editing at approximately 30 FPS on a single Nvidia B200 GPU. Code is available at https://github.com/jd-opensource/JoyAI-Video-Edit.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.