Ripple: 크로스 모달 순환 메모리를 이용한 실시간 오디오-비디오 생성
Ripple: Real-Time Streaming Audio-Video Generation With Cross-Modal Recurrent Memory
오디오-비디오 생성 모델은 뛰어난 품질을 달성하지만, 높은 지연 시간으로 인해 실시간 애플리케이션에 적합하지 않습니다. 몇 가지 스트리밍 오디오-비디오 생성 방법이 제안되었지만, 여전히 비용이 많이 들고 장편 생성이 어렵습니다. 이러한 문제를 해결하기 위해, 우리는 크로스 모달 순환 메모리 메커니즘을 갖춘 실시간 통합 오디오-비디오 생성 시스템인 **Ripple**을 제안합니다. Ripple은 효율적인 스트리밍 추론과 함께 장기적인 맥락을 유지하기 위해, 고정 길이 슬라이딩 윈도우 어텐션과 모달리티별 메모리 상태를 결합하여 오디오 및 비디오 컨텍스트를 지속적으로 요약합니다. 또한, 오디오-비주얼 동기화를 향상시키기 위해 크로스 모달 메모리 상호 작용을 도입했습니다. 이 메모리 강화 모델을 효과적으로 학습하기 위해, 우리는 세 단계의 학습 방법을 고안했습니다: (1) 양방향 오디오-비디오 튜처를 사용하여 시뮬레이션된 메모리가 있는 블록 단위 인과 관계 어텐션으로 적응시키고, (2) 엔드투엔드 증류를 통해 메모리 구성 및 상호 작용 파이프라인을 최적화하고, (3) 스트리밍 오디오-비디오 생성에 특화된 온라인 강화 후학습을 적용합니다. 그 결과, Ripple은 480P 해상도에서 약 28 FPS의 프레임 속도를 달성하며, 이는 튜처보다 빠르며 일관성 있는 장편 생성이 가능합니다. 짧은 비디오와 긴 비디오 벤치마크에 대한 광범위한 실험 결과는 기존 오프라인 및 온라인 통합 오디오-비디오 생성 방법에 비해 우수한 성능을 보여줍니다.
Audio-video generative models achieve impressive quality but suffer from high latency, making them unsuitable for real-time applications. Although several streaming audio-video generation methods have been proposed, they remain costly and fail to support long-form generation. To address this, we propose \textbf{Ripple}, a real-time joint audio-video generation system with a cross-modal recurrent memory mechanism. To enable efficient streaming inference while preserving long-term context, Ripple combines a fixed-length sliding-window attention with modality-specific memory states that continuously summarize audio and video context. Cross-modal memory interaction is further introduced to enhance audio-visual synchronization. To learn this memory-augmented model effectively, we devise a three-stage training recipe: (1) adapting a bidirectional audio-video teacher to block-wise causal attention with simulated memory, (2) optimizing the memory construction and interaction pipeline through end-to-end distillation, and (3) applying online reinforcement post-training tailored for streaming audio-video generation. As a result, Ripple achieves ~28 FPS at 480P resolution, over faster than the teacher, while capable of coherent long-form generation. Extensive experiments on both short-video and long-video benchmarks demonstrate our superior performance over existing offline and online joint audio-video generation methods.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.