DiFlowDubber: 교차 모달 정렬 및 동기화를 통한 자동 비디오 더빙을 위한 이산 흐름 매칭
DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization
비디오 더빙은 영화 제작, 멀티미디어 콘텐츠 제작, 보조 음성 기술 등 다양한 분야에서 활용됩니다. 기존 방법들은 제한된 더빙 데이터셋에 직접 학습하거나, 사전 학습된 텍스트 음성 변환(TTS) 모델을 사용하는 2단계 파이프라인을 채택하는 경우가 많습니다. 하지만 이러한 방법들은 종종 표현력이 풍부한 음성, 다채로운 음향 특징, 그리고 정확한 동기화를 구현하는 데 어려움을 겪습니다. 이러한 문제점을 해결하기 위해, 우리는 이산 흐름 매칭 생성 모델을 기반으로 하는 새로운 2단계 학습 프레임워크인 DiFlowDubber를 제안합니다. 구체적으로, 우리는 얼굴 표정을 통해 전반적인 음성 억양과 스타일 단서를 포착하는 FaPro 모듈을 설계하고, 이를 통해 후속 음성 속성 모델링을 안내합니다. 또한, 텍스트, 비디오, 음성 간의 모달리티 격차를 해소하고, 음성-입술 동기화를 보장하기 위해 Synchronizer 모듈을 도입하여 교차 모달 정렬을 개선하고, 입술 움직임과 시간적으로 동기화된 음성을 생성합니다. 두 개의 주요 벤치마크 데이터셋에 대한 실험 결과, DiFlowDubber는 여러 지표에서 기존 방법들보다 우수한 성능을 보였습니다.
Video dubbing has broad applications in filmmaking, multimedia creation, and assistive speech technology. Existing approaches either train directly on limited dubbing datasets or adopt a two-stage pipeline that adapts pre-trained text-to-speech (TTS) models, which often struggle to produce expressive prosody, rich acoustic characteristics, and precise synchronization. To address these issues, we propose DiFlowDubber with a novel two-stage training framework that effectively transfers knowledge from a pre-trained TTS model to video-driven dubbing, with a discrete flow matching generative backbone. Specifically, we design a FaPro module that captures global prosody and stylistic cues from facial expressions and leverages this information to guide the modeling of subsequent speech attributes. To ensure precise speech-lip synchronization, we introduce a Synchronizer module that bridges the modality gap among text, video, and speech, thereby improving cross-modal alignment and generating speech that is temporally synchronized with lip movements. Experiments on two primary benchmark datasets demonstrate that DiFlowDubber outperforms previous methods across multiple metrics.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.