2606.26534v1 Jun 25, 2026 cs.SD

VoiceTTA: 강화 학습 기반 테스트 시간 적응을 통한 제로샷 텍스트 음성 변환 성능 향상

VoiceTTA: Enhancing Zero-Shot Text-to-Speech via Reinforcement Learning-Based Test-Time Adaptation

Tianxin Xie
Tianxin Xie
Citations: 90
h-index: 5
Li Liu
Li Liu
Citations: 3,394
h-index: 8
Chenxing Li
Chenxing Li
Citations: 156
h-index: 6
Dong Yu
Dong Yu
Citations: 233
h-index: 7

최근, 제로샷 텍스트 음성 변환(TTS) 기술은 고품질의 표현력 있는 음성 합성을 가능하게 하지만, 흔하지 않은 상황에서의 새로운 발화 스타일(예: 토크쇼, 방언)을 모방하는 데 어려움을 겪는 경우가 많습니다. 또한, 사전 학습된 모델을 미세 조정하려면 대규모의 고품질 데이터셋이 필요하며, 이는 빠른 개인화 과정을 제한합니다. 본 논문에서는 강화 학습 기반 테스트 시간 적응(TTA) 방법인 VoiceTTA를 제안하여, 사전 학습된 제로샷 TTS 모델의 음성 모방 성능을 향상시킵니다. VoiceTTA는 F0 및 에너지 값의 변동 차이를 기반으로 한 두 가지 스타일 보상과 함께, 화자 유사성 및 가독성(사전 학습된 Whisper 모델로부터 얻은 WER)을 활용하며, 플로우 매칭 기반 모델에서 그룹 상대 선호도 최적화(GRPO)를 통해 추론 시에 학습 가능한 접두사를 최적화합니다. 광범위한 실험 결과는 VoiceTTA가 흔하지 않은 음성 프롬프트에서 상당한 성능 향상을 보여주며, 최첨단 기준 모델보다 우수한 성능을 발휘함을 입증했습니다. 오디오 샘플은 https://voicetta.pages.dev/ 에서 확인할 수 있습니다.

Original Abstract

Recently, zero-shot text-to-speech (TTS) has enabled high-fidelity and expressive speech synthesis, but it often fails to imitate unseen speaking styles from uncommon scenarios (e.g., crosstalk, dialects). Moreover, fine-tuning pretrained models requires large, high-quality datasets, limiting rapid personalization. We propose VoiceTTA, a reinforcement learning-based test-time adaptation (TTA) method that improves voice imitation of pretrained zero-shot TTS models. VoiceTTA introduces two style rewards based on coefficient-of-variation differences of F0 and energy, combined with speaker similarity and intelligibility (WER from a pretrained Whisper model), and optimizes learnable prefixes via group relative preference optimization (GRPO) in a flow matching-based model at inference time. Extensive experiments demonstrate substantial improvements on uncommon speech prompts, outperforming state-of-the-art baselines. Audio samples are available at https://voicetta.pages.dev/

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!