변조(Spoofing) 방지 기능이 있는 화자 인증을 위한 대규모 오디오 언어 모델
Large Audio Language Models for Spoofing-Aware Speaker Verification
최근 텍스트 음성 변환 및 음성 복제 기술의 발전으로 인해 고품질의 변조 공격이 저렴하고 확장 가능해짐에 따라, 특히 자동화된 화자 인증 시스템(ASV)에 심각한 위협을 가하고 있습니다. 기존의 방어 방법은 주로 딥페이크 탐지를 위한 이진 대응책(CM) 또는 변조 방지 화자 인증(SASV)을 통해 이러한 위협에 대처하며, 현재 시스템은 모듈식 ASV-CM 융합 및 연결된 파이프라인으로 구성됩니다. 대규모 오디오 언어 모델(LALM)은 CM 및 ASV를 포함한 관련 오디오 작업에서 유망한 결과를 보여주었지만, 검증 가능성과 판별적 예측을 넘어 자연스러운 언어로 설명하는 능력을 가진 LALM의 SASV 적용은 아직 탐구되지 않았습니다. 본 연구에서는 다양한 방법을 통해 LALM을 SASV에 적용하고 기존 파이프라인과 비교 평가합니다. 여기에는 제로샷 프롬프트, 지도 학습 기반 적응, 추론 중심 훈련 및 강화 학습 기반 최적화가 포함됩니다. 실험 결과는 사전 훈련된 LALM이 제로샷 환경에서는 우수한 성능을 보이지 않지만(무작위 수준), 특정 작업에 맞게 모델을 조정하면 이러한 격차를 줄일 수 있음을 보여줍니다. 또한, 여러 가지 서로 다른 방식으로 경쟁력 있는 SASV 성능을 달성할 수 있다는 것을 확인했습니다. 이러한 결과는 LALM을 통합된 SASV 시스템의 유망하고 검증 가능한 기반 기술로 제시하며, 동시에 기존 연결 방식이 여전히 우수한 영역은 무엇인지 명확히 합니다.
Recent advances in text-to-speech and voice cloning make high-quality spoofing inexpensive and scalable, threatening voice authentication systems, especially automatic speaker verification (ASV). Existing defenses mainly address this threat through binary countermeasures (CMs) for deepfake detection or spoofing-aware speaker verification (SASV), where current systems are dominated by modular ASV-CM fusion and cascaded pipelines. Although large audio language models (LALMs) have shown promise on related audio tasks, including CM and ASV, their use for SASV remains unexplored, despite their capacity to produce natural-language rationales for auditing and robustness beyond discriminative predictions. This work systematically evaluates LALMs for SASV against conventional pipelines under zero-shot prompting, supervised adaptation, reasoning-oriented training, and reinforcement-learning-based optimization. Our results show that pretrained LALMs are near chance in the zero-shot setting, confirming that they are not natively suited to SASV, but that task-specific adaptation closes this gap. We further find that competitive SASV performance can be achieved through several distinct routes. These findings position LALMs as a promising and auditable foundation for unified SASV, while clarifying where conventional cascade systems still lead.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.