2608.03974v1 Aug 04, 2026 cs.CV

JoyAI-Video-Edit: 자기 회귀 확산 모델을 활용한 실시간, 개방형 비디오 편집

JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion

Nan Duan
Nan Duan
Citations: 43
h-index: 3
Yuan Zhang
Yuan Zhang
Citations: 27
h-index: 2
Guoqing Ma
Guoqing Ma
Citations: 216
h-index: 4
Haoyang Huang
Haoyang Huang
Citations: 397
h-index: 5
Hang Xu
Hang Xu
Citations: 15
h-index: 2
Jie Huang
Jie Huang
Citations: 41
h-index: 3
Yitong Li
Yitong Li
Citations: 14
h-index: 2
Xinran Qin
Xinran Qin
Citations: 10
h-index: 2
Guohui Zhang
Guohui Zhang
Citations: 44
h-index: 4
Siming Fu
Siming Fu
Citations: 343
h-index: 8
Lin Song
Lin Song
Citations: 1,726
h-index: 10
Yukang Chen
Yukang Chen
Citations: 6,183
h-index: 31
Yicheng Xiao
Yicheng Xiao
Citations: 186
h-index: 7
Wenxun Dai
Wenxun Dai
Citations: 64
h-index: 2
Jian-min Yuan
Jian-min Yuan
Citations: 81
h-index: 5
Xiaojuan Qi
Xiaojuan Qi
Citations: 170
h-index: 5
Tommy Zhang
Tommy Zhang
Citations: 0
h-index: 0
Wenbo Li
Wenbo Li
Citations: 249
h-index: 4
Wei Huang
Wei Huang
Citations: 42
h-index: 2
Chuyang Zhao
Chuyang Zhao
Citations: 141
h-index: 4
Peihao Li
Peihao Li
Citations: 255
h-index: 5
Maoquan Zhang
Maoquan Zhang
Citations: 2
h-index: 1
Xuying Zhang
Xuying Zhang
Citations: 0
h-index: 0
Xin Han
Xin Han
Citations: 189
h-index: 3
Shuai Lu
Shuai Lu
Citations: 7,212
h-index: 14

실시간 비디오 편집은 낮은 지연 시간으로 원본의 충실도를 유지하고 장기적인 시간적 일관성을 확보하면서 제한된 계산 자원을 사용해야 합니다. 본 논문에서는 미래 프레임이나 미리 정의된 비디오 길이에 의존하지 않고 실시간, 개방형 비디오 편집을 가능하게 하는 160억 개의 파라미터를 가진 자기 회귀 확산 모델 기반 프레임워크인 JoyAI-Video-Edit를 제안합니다. 저희 방법은 청크별 자기 회귀 적응, 소스 기반 분포 일치 증류(SA-DMD), 그리고 장기 자기 회귀 증류를 결합하여 학습과 추론 간의 불일치를 줄이고, 두 단계 생성 과정에서 원본의 충실도를 유지하며, 누적된 시간적 드리프트를 완화합니다. 광범위한 자동 및 인간 평가 결과, JoyAI-Video-Edit는 기존 스트리밍 편집 도구를 크게 능가하며, 짧은 비디오와 긴 비디오 모두에서 강력한 오프라인 시스템과 경쟁력 있는 성능을 보여줍니다. 전체 시스템은 단일 Nvidia B200 GPU에서 약 30 FPS의 속도로 720p 비디오 편집을 수행할 수 있습니다. 코드는 https://github.com/jd-opensource/JoyAI-Video-Edit 에서 확인할 수 있습니다.

Original Abstract

Real-time video editing requires low-latency causal generation with bounded computational resources while preserving source fidelity and long-term temporal consistency. We present JoyAI-Video-Edit, a 16B-parameter autoregressive diffusion framework for real-time, open-ended video editing without access to future frames or a predefined video duration. Our method combines chunk-wise autoregressive adaptation, Source-Anchored Distribution Matching Distillation (SA-DMD), and Long-Horizon Autoregressive Distillation to reduce train--inference mismatch, preserve source fidelity during two-step generation, and mitigate accumulated temporal drift. Extensive automatic and human evaluations show that JoyAI-Video-Edit substantially outperforms existing streaming editors and remains competitive with strong offline systems on both short and long videos. The complete system achieves end-to-end 720p video editing at approximately 30 FPS on a single Nvidia B200 GPU. Code is available at https://github.com/jd-opensource/JoyAI-Video-Edit.

1 Citations
0 Influential
0 Altmetric
6.9 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!