재구성 왜곡에 대한 강건성을 위한 특징 정렬 음성 워터마킹
Feature-Aligned Speech Watermarking for Robustness to Reconstruction Distortions
음성 워터마킹은 식별 가능한 정보를 오디오에 삽입하되, 인간의 청각으로 감지되지 않도록 하는 기술입니다. 기존 방법들은 원본 오디오의 지각적 품질을 유지하기 위해 고충실도, 저에너지 설계를 채택하지만, 결과적으로 생성된 워터마크는 음성 재구성 모델에 의한 제거 시 강건성이 부족합니다. 기존 설계 방식에서 존재하는 강건성-충실도 간의 상충 관계 때문에 강건성을 향상시키는 것은 어렵습니다. 왜냐하면 워터마크 에너지를 증가시키면 강건성은 향상되지만 충실도가 감소하기 때문입니다. 이러한 문제를 해결하기 위해, 본 연구에서는 워터마크를 원본 음성 특징 분포와 정렬하는 특징 기반 워터마킹 방법을 제안합니다. 이를 통해 더 높은 워터마크 에너지를 사용하여 강건성을 향상시키면서도 지각적 투명성을 유지할 수 있습니다. 사전 학습된 음성 코덱을 사용하여 가짜 음성 워터마크를 생성하고, 입력 오디오의 스펙트로그램에 통합하며, VAD 손실 및 지각적 손실을 활용하여 음성 구간 내에서 워터마크를 삽입합니다. 실험 결과는 제안하는 방법이 기존 방식과 유사한 수준의 지각적 투명성을 유지하면서, 알려진 및 알려지지 않은 음성 재구성 모델 하에서도 현저히 향상된 강건성을 제공함을 보여줍니다.
Audio watermarking aims to embed identifiable information into audio while remaining imperceptible. Existing methods adopt high-fidelity, low-energy designs to preserve perceptual quality, but the resulting watermarks lack robustness under suppression by speech reconstruction models. Improving robustness is challenging due to the inherent robustness-fidelity trade-off in existing designs, where increasing watermark energy improves robustness but reduces fidelity. To address this problem, we propose a feature-aligned watermarking method that aligns the watermark with the original speech feature distribution, allowing higher watermark energy to improve robustness while preserving imperceptibility. We use a pretrained speech codec to generate a pseudo-speech watermark and fuse it into the spectrogram of the input audio, with VAD loss and perceptual losses guiding embedding within voiced regions. Experiments show that our method maintains imperceptibility comparable to existing approaches while substantially improving robustness under both seen and unseen speech reconstruction models.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.