말하는 것에서 노래하는 것으로: 오디오-비주얼 딥페이크 탐지를 위한 새로운 도전
From Talking to Singing: A New Challenge for Audio-Visual Deepfake Detection
오디오-비주얼 생성 모델의 급속한 발전으로 인해 신뢰할 수 있는 위조 탐지는 점점 더 중요해지고 있습니다. 기존의 오디오-비주얼 딥페이크 탐지 방법은 일반적으로 상호 모달 불일치를 기반으로 합니다. 노래에서는 리듬적인 발성이 이러한 연관성을 약화시키고, 상당한 도메인 변화를 야기하여 탐지 성능을 크게 저하시킵니다. 우리는 리듬 인식을 갖춘 생성 모델을 사용하여 노래 벤치마크의 격차를 해소하기 위해 Singing Head DeepFake (SHDF) 데이터 세트를 구축했습니다. 다양한 시나리오에서의 도메인 변화에 대처하기 위해, 말하는 것과 노래하는 것 모두에서 일반화 가능한 Text-guided Audio-Visual Forgery Detection (T-AVFD) 프레임워크를 제안합니다. T-AVFD는 얼굴 진위 패턴 학습기 및 다중 모달 차등 가중치 학습 모듈로 구성됩니다. 패턴 학습기는 얼굴 특징을 다양한 수준의 텍스트 설명과 연결하여 일반화 가능한 진위 패턴을 학습합니다. 가중치 학습 모듈은 고유한 오디오-비주얼 일관성을 유지하고, 차등 가중치를 통해 이를 진위 패턴과 적응적으로 통합합니다. 여러 말하는 얼굴 딥페이크 데이터 세트와 SHDF에 대한 광범위한 실험 결과, 기존의 기본 모델보다 일관된 성능 향상과 다양한 변동 조건 하에서의 강력한 견고성을 보여줍니다.
With rapid advances in audio-visual generative models, reliable forgery detection becomes increasingly critical. Existing methods for audio-visual deepfake detection typically rely on cross-modal inconsistencies. In singing, rhythmic vocalization weakens this coupling and introduces a nontrivial domain shift, substantially degrading detection performance. We construct the Singing Head DeepFake (SHDF) dataset using rhythm-aware generative models to fill the gap in singing benchmarks. To cope with cross-scenario domain shifts, we propose a Text-guided Audio-Visual Forgery Detection (T-AVFD) framework that generalizes across both talking and singing scenarios. T-AVFD comprises a facial authenticity pattern learner and a multi-modal differential weight learning module. The pattern learner aligns facial features with multi-granularity textual descriptions to learn generalizable authenticity patterns. The weight learning module preserves intrinsic audio-visual consistency and adaptively integrates it with authenticity patterns via differential weighting. Extensive experiments on multiple talking head deepfake datasets and SHDF show consistent improvements over existing baselines and strong robustness under diverse perturbations.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.