일반화 가능한 오디오 딥페이크 탐지를 위한 이중 수준 직교 분리 학습
Dual-Granularity Orthogonal Disentanglement for Generalizable Audio Deepfake Detection
오디오 딥페이크 탐지기는 종종 화자 정보를 학습하여 합성된 특징이 아닌 화자 식별 정보에 의존하기 때문에 화자에 따라 성능이 달라지는 경향이 있습니다. 이러한 문제를 해결하기 위한 기존 방법들은 복잡한 구조를 가지거나 학습 과정에서 불안정성을 야기합니다. 본 논문에서는 두 가지 수준에서 특징의 독립성을 강화하는 이중 수준 직교 분리 학습 프레임워크를 제안합니다. 샘플 레벨에서의 코사인 직교성은 방향성 상관관계를 줄이고, 배치 레벨에서의 교차 공분산 정규화는 임베딩 차원 간의 선형 상관관계를 제거합니다. 제안된 방법은 보조 네트워크나 적대적 학습 없이 점진적으로 직교성 제약을 강화하는 학습 스케줄을 사용합니다. ASVspoof 2019 LA, ASVspoof 2021 DF, 그리고 실제 환경 데이터 세트에서의 실험 결과는 제안된 방법이 각각 1.35%, 7.88% 및 21.58%의 동일 오류율(EER)을 달성했으며, 특히 다른 데이터 세트로 성능 평가 시 gradient reversal disentanglement 방법보다 2.60%p 더 높은 성능을 보였습니다.
Audio deepfake detectors often fail to generalize across speakers, as they learn speaker-identity features rather than synthesis artifacts, known as implicit identity leakage. Existing methods address this but incur architectural complexity or training instability. This paper proposes a dual-granularity orthogonal disentanglement framework enforcing feature independence at two levels: sample-level cosine orthogonality captures directional decorrelation, while batch-level cross-covariance regularization eliminates linear correlations across embedding dimensions. A curriculum disentanglement schedule progressively strengthens the orthogonality constraint without auxiliary networks or adversarial dynamics. Experiments on ASVspoof 2019 LA, ASVspoof 2021 DF, and In-the-Wild datasets demonstrate that the proposed method achieves 1.35%, 7.88%, and 21.58% equal error rates (EER), respectively, surpassing gradient reversal disentanglement by 2.60% absolute on cross-dataset transfer.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.