Vorch-Streamer: 실시간 장문 스트리밍을 위한 인간의 오디오-비주얼 생성 확장
Vorch-Streamer: Extending Human Audio-Visual Generation to Real-Time Long-Form Streaming
실시간으로 긴 형식의 아바타 오디오-비디오 생성을 위해서는 인과 관계를 가지는 연속적인 합성이 필요하며, 동시에 오디오와 비디오의 동기화 및 시각적 일관성을 유지해야 합니다. 사전 훈련된 양방향 모델을 이러한 환경에 적용하는 것은 두 가지 주요 난제를 야기합니다. 첫째, 생성된 블록을 재사용하여 문맥 정보를 제공하면 노출 편향이 발생하여 오류가 누적되고 시간이 지남에 따라 시각적인 왜곡이 나타날 수 있습니다. 둘째, 전체 음성 발화는 원인-결과 관계를 명확하게 제시하지 않으므로, 제한된 로컬 오디오-비디오 문맥 정보만 주어졌을 때 어떤 부분을 다음에 말해야 하는지 결정하기 어렵습니다. 본 논문에서는 이러한 과제를 해결하고 실시간 장문 텍스트-오디오-비디오(T2AV) 스트리밍을 가능하게 하는 후처리 프레임워크인 **Vorch-Streamer**를 제시합니다. 우리는 12~21초 길이의 아바타 클립으로 구성된 80,000개의 합성 데이터셋을 구축하고, 먼저 가이드 포싱과 디퓨전 포싱을 혼합하여 사용한 인과적 생성기를 학습했습니다. 그런 다음, DMD 증류를 활용한 장기 자기 포싱(long-horizon Self Forcing)을 적용하여 모델이 자신의 출력 분포에 노출되도록 하면서, 사전 훈련된 양방향 모델의 품질을 유지했습니다. 음성 진행을 명시적으로 제어하기 위해, 외부 언어 모델이 이산적인 25Hz 음성 계획 토큰을 예측하며, 이러한 토큰의 연속적인 특징은 오디오 디퓨전 브랜치에 영향을 미쳐 각 인과적 블록이 발음해야 할 내용과 일치하도록 합니다. 제한된 인과 관계 문맥 정보와 4단계 노이즈 제거를 통해 Vorch-Streamer는 27.12 FPS의 속도로 텍스트로부터 오디오와 비디오를 동시에 생성하며, 이는 24 FPS의 실시간 재생 속도를 능가합니다. 또한, Vorch-Streamer는 우수한 오디오-입술 동기화 및 장기간 생성 과정에서 강한 동일성 유지 기능을 제공합니다.
Real-time long-form avatar audio-video generation requires causal, continuous synthesis while maintaining audiovisual synchronization and visual consistency. Adapting a pretrained bidirectional model to this setting presents two key dilemmas. First, autoregressively reusing generated blocks as context creates exposure bias, causing errors and visual drift to accumulate over long rollouts. Second, a global speech utterance does not indicates a causal generator which portion should be spoken next when only limited local audio-video context is available. We present Vorch-Streamer, a post-training framework that addresses these challenges and enables real-time long-form Text-to-Audio-Video (T2AV) streaming. We construct a synthetic corpus of 80K avatar clips spanning 12-21 seconds and first train a causal generator with mixed Teacher Forcing and Diffusion Forcing. We then apply long-horizon Self Forcing with DMD distillation, exposing the model to its own rollout distribution while preserving the quality of the pretrained bidirectional teacher. To explicitly control speech progression, an external language model predicts discrete 25-Hz speech-planning tokens, whose continuous features condition the audio diffusion branch and align each causal block with the content it should speak. With bounded causal context and four-step denoising, Vorch-Streamer jointly generates audio and video from text at 27.12 FPS, exceeding the 24-FPS real-time playback rate while maintaining competitive audio-lip synchronization and strong identity preservation over long-form generation.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.