ZeSTA: 도메인 조건부 훈련을 통한 제로샷 TTS 증강 기법을 활용한 데이터 효율적인 개인 맞춤형 음성 합성
ZeSTA: Zero-Shot TTS Augmentation with Domain-Conditioned Training for Data-Efficient Personalized Speech Synthesis
본 연구에서는 데이터가 부족한 개인 맞춤형 음성 합성을 위한 데이터 증강 방법으로 제로샷 텍스트-음성 변환(ZS-TTS)의 활용 가능성을 탐구합니다. 합성 데이터 증강은 언어적으로 풍부하고 음성적으로 다양한 음성을 제공할 수 있지만, 제한된 실제 녹음 데이터와 대량의 합성 음성을 무분별하게 혼합하면 미세 조정 과정에서 화자 유사성이 저하되는 문제가 발생할 수 있습니다. 이러한 문제를 해결하기 위해, 우리는 경량화된 도메인 임베딩을 사용하여 실제 음성과 합성 음성을 구분하고, 실제 데이터의 과대 표본 추출을 통해 극히 제한적인 대상 데이터 환경에서도 안정적인 적응을 가능하게 하는 간단한 도메인 조건부 훈련 프레임워크인 ZeSTA를 제안합니다. LibriTTS 데이터셋과 자체 제작 데이터셋에서 두 가지 ZS-TTS 소스를 활용한 실험 결과, 제안하는 방법은 기존의 합성 데이터 증강 방식보다 화자 유사성을 향상시키면서도 음성 명료도와 지각적 품질을 유지하는 것을 확인했습니다.
We investigate the use of zero-shot text-to-speech (ZS-TTS) as a data augmentation source for low-resource personalized speech synthesis. While synthetic augmentation can provide linguistically rich and phonetically diverse speech, naively mixing large amounts of synthetic speech with limited real recordings often leads to speaker similarity degradation during fine-tuning. To address this issue, we propose ZeSTA, a simple domain-conditioned training framework that distinguishes real and synthetic speech via a lightweight domain embedding, combined with real-data oversampling to stabilize adaptation under extremely limited target data, without modifying the base architecture. Experiments on LibriTTS and an in-house dataset with two ZS-TTS sources demonstrate that our approach improves speaker similarity over naive synthetic augmentation while preserving intelligibility and perceptual quality.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.