2607.26818v1 Jul 29, 2026 cs.CV

Ripple: 크로스 모달 순환 메모리를 이용한 실시간 오디오-비디오 생성

Ripple: Real-Time Streaming Audio-Video Generation With Cross-Modal Recurrent Memory

Zhixiang He
Zhixiang He
Citations: 56
h-index: 2
Yanbo Ding
Yanbo Ding
Citations: 35
h-index: 3
Zhizhi Guo
Zhizhi Guo
Citations: 9
h-index: 1
Quanyue Song
Quanyue Song
Citations: 2
h-index: 1
Yishan He
Yishan He
Citations: 1
h-index: 1
Yongxiang Li
Yongxiang Li
Citations: 0
h-index: 0
Yali Wang
Yali Wang
Citations: 4
h-index: 1

오디오-비디오 생성 모델은 뛰어난 품질을 달성하지만, 높은 지연 시간으로 인해 실시간 애플리케이션에 적합하지 않습니다. 몇 가지 스트리밍 오디오-비디오 생성 방법이 제안되었지만, 여전히 비용이 많이 들고 장편 생성이 어렵습니다. 이러한 문제를 해결하기 위해, 우리는 크로스 모달 순환 메모리 메커니즘을 갖춘 실시간 통합 오디오-비디오 생성 시스템인 **Ripple**을 제안합니다. Ripple은 효율적인 스트리밍 추론과 함께 장기적인 맥락을 유지하기 위해, 고정 길이 슬라이딩 윈도우 어텐션과 모달리티별 메모리 상태를 결합하여 오디오 및 비디오 컨텍스트를 지속적으로 요약합니다. 또한, 오디오-비주얼 동기화를 향상시키기 위해 크로스 모달 메모리 상호 작용을 도입했습니다. 이 메모리 강화 모델을 효과적으로 학습하기 위해, 우리는 세 단계의 학습 방법을 고안했습니다: (1) 양방향 오디오-비디오 튜처를 사용하여 시뮬레이션된 메모리가 있는 블록 단위 인과 관계 어텐션으로 적응시키고, (2) 엔드투엔드 증류를 통해 메모리 구성 및 상호 작용 파이프라인을 최적화하고, (3) 스트리밍 오디오-비디오 생성에 특화된 온라인 강화 후학습을 적용합니다. 그 결과, Ripple은 480P 해상도에서 약 28 FPS의 프레임 속도를 달성하며, 이는 튜처보다 빠르며 일관성 있는 장편 생성이 가능합니다. 짧은 비디오와 긴 비디오 벤치마크에 대한 광범위한 실험 결과는 기존 오프라인 및 온라인 통합 오디오-비디오 생성 방법에 비해 우수한 성능을 보여줍니다.

Original Abstract

Audio-video generative models achieve impressive quality but suffer from high latency, making them unsuitable for real-time applications. Although several streaming audio-video generation methods have been proposed, they remain costly and fail to support long-form generation. To address this, we propose \textbf{Ripple}, a real-time joint audio-video generation system with a cross-modal recurrent memory mechanism. To enable efficient streaming inference while preserving long-term context, Ripple combines a fixed-length sliding-window attention with modality-specific memory states that continuously summarize audio and video context. Cross-modal memory interaction is further introduced to enhance audio-visual synchronization. To learn this memory-augmented model effectively, we devise a three-stage training recipe: (1) adapting a bidirectional audio-video teacher to block-wise causal attention with simulated memory, (2) optimizing the memory construction and interaction pipeline through end-to-end distillation, and (3) applying online reinforcement post-training tailored for streaming audio-video generation. As a result, Ripple achieves ~28 FPS at 480P resolution, over faster than the teacher, while capable of coherent long-form generation. Extensive experiments on both short-video and long-video benchmarks demonstrate our superior performance over existing offline and online joint audio-video generation methods.

0 Citations
0 Influential
1.5 Altmetric
7.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!