파형 강건성 너머: 자동 음성 인식 시스템에 대한 강력한 특징-보코더 적대적 공격
Beyond Waveform Robustness: Robust Feature-Vocoder Adversarial Attacks on Automatic Speech Recognition
자동 음성 인식(ASR) 시스템은 다국어 음성-텍스트 변환에 널리 사용되고 있으며, 이러한 시스템의 적대적 공격에 대한 강건성은 중요한 연구 주제입니다. 기존의 적대적 공격은 주로 음성 오디오에 직접적으로 적대적인 노이즈를 추가하는 방식으로 작동합니다. 그러나 이전 연구에서 기존의 적대적 공격은 두 가지 한계를 가지고 있음이 밝혀졌습니다. 즉, 블랙박스 ASR 시스템으로 잘 전이되지 않으며, 입력 공간의 변화에 특화된 방어 기법에 의해 점점 더 완화되고 있습니다. 본 연구에서는 클린 참조 특징-보코더 공격(Clean-Referenced Feature-Vocoder Attack)이라는 새로운 접근 방식을 제안합니다. 이는 서브 모델 기반의 블랙박스 공격으로, 적대적인 탐색 공간을 원시 파형에서 자기 지도 학습(SSL) 표현으로 이동시킵니다. 전이성 문제를 해결하기 위해, 우리는 파형 샘플과 같은 저수준 정보 대신 더 일반화된 음향-음운적 표현을 변경합니다. 이를 통해 서브 모델에 특정한 파형 기울기에 대한 의존성을 줄이고 ASR 시스템 전체에서 일반화되는 적대적인 변화를 유도합니다. 또한 다양한 방어 기법을 우회하기 위해, 우리는 명시적인 추가 파형 노이즈 대신 SSL 특징 공간의 변화를 활용하고 보코더를 사용하여 이를 음성 형태의 파형 적대적 신호로 재구성합니다. 이렇게 하면 결과 샘플이 파형 기반 방어에 덜 민감하게 됩니다. 광범위한 실험 결과, 우리의 공격은 공개 서브 모델인 Whisper-small에서만 최적화되었음에도 불구하고 블랙박스 ASR 모델로 효과적으로 전이되어 최고 성능(SOTA) 기준보다 +26.6%의 단어 오류율(WER) 개선을 달성했습니다. 또한 여러 가지 학습 기반 방어 기법에 대해서도 +36.2%의 WER 개선을 보여주며, 이는 현재 ASR 시스템의 강건성 평가에 존재하는 중요한 문제점을 드러냅니다.
Automatic speech recognition (ASR) systems have become widely used for multilingual speech-to-text transcription. Their robustness to adversarial attacks has become an important topic for the community. Existing adversarial attacks directly add adversarial noise to the speech audio. However, prior work has shown that existing adversarial attacks face two limitations: they often transfer poorly to black-box ASR systems and are increasingly mitigated by defenses tailored to input-space perturbations. In this work, we propose a Clean-Referenced Feature-Vocoder Attack, a surrogate-based black-box attack that moves the adversarial search space from raw waveforms to self-supervised learning (SSL) representations. To address the transferability limitation, we perturb more generalizable acoustic-phonetic representations rather than low-level waveform samples, reducing dependence on surrogate-specific waveform gradients and encouraging adversarial perturbations that generalize across ASR systems. To bypass different defenses, we shift the adversarial signal from explicit additive waveform noise to SSL feature-space perturbations and reconstruct them through a vocoder into speech-like waveform adversarial signals, making the resulting samples less aligned with waveform-bounded defenses. Extensive experiments show that, when optimized only on raw Whisper-small as a public surrogate model, our attack transfers effectively to black-box ASR models with a +26.6 WER improvement over the SOTA baseline, while also remaining effective against multiple training defenses with a +36.2 WER improvement. These results reveal a blind spot in current ASR robustness evaluation.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.