2608.03215v1 Aug 04, 2026 eess.AS

GROW: 그룹 상대적 장점 가중 온폴리시 강화 학습을 이용한 자기회귀-확산 기반 텍스트 음성 변환 모델

GROW: Group-Relative Advantage-Weighted On-Policy Reinforcement Learning of Autoregressive-Diffusion Text-to-Speech model

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Tian Tan
Tian Tan
Citations: 863
h-index: 6
Zhikang Niu
Zhikang Niu
Citations: 831
h-index: 10
Xie Chen
Xie Chen
Citations: 16
h-index: 2
Qian Chen
Qian Chen
Citations: 767
h-index: 8
Ya-Zhen Song
Ya-Zhen Song
Citations: 257
h-index: 7
Wenming Tu
Wenming Tu
Citations: 20
h-index: 2

플로우 매칭 텍스트 음성 변환에서 강화 학습은 결정론적인 ODE 샘플링으로 인해 복잡해집니다. 기존의 경로 수준 정책 경사 방법은 일반적으로 ODE를 SDE로 변환하고 단계별 가능도 비율을 추적하며, 이는 확률적 교란과 상당한 오버헤드를 발생시킵니다. 본 연구에서는 표준 플로우 매칭 목표에 직접 작용하는 그룹 상대적 장점 가중 온폴리시 강화 학습 방법인 GROW를 제안합니다. 각 프롬프트에 대해 GROW는 온폴리시 발화를 그룹으로 샘플링하고, 각 그룹 내에서 음성 명료도 및 화자 유사도 보상을 개별적으로 표준화한 후 이를 결합하여 플로우 매칭 회귀를 재가중합니다. Wasserstein-2 속도 페널티는 업데이트된 모델을 사전 훈련된 참조 모델에 고정시키는 역할을 합니다. 그룹 평균 보상 기준선이 도입되어 보상 가중치를 장점 가중치로 변환합니다. 강력한 사전 훈련된 TTS 모델에서, 집중된 보상은 보상과 무관한 자기 모방으로 이어지는 반면, 평균이 0인 서명된 장점을 사용하면 그룹 내에서 효과적인 신용 할당을 유지할 수 있습니다. DiTAR에 구현된 GROW는 LibriSpeech 및 Seed-TTS EN/ZH 데이터셋에서 WER(단어 오류율)을 2.016에서 1.558로 줄이고, 화자 유사도를 0.676에서 0.715로 향상시키면서 UTMOS(Unintended Text-to-Music Overlay Score)를 유지합니다. 10-NFE(Non-flow Evaluation) 훈련과 32-NFE 평가를 수행한 GROW는 32-NFE DiTAR-GRPO에 비해 2.9배 더 빠르게 학습하면서도 유사한 성능을 유지합니다. 우리는 GROW의 전체 코드, 정확한 DiTAR 재현 결과 및 모든 모델 체크포인트를 공개할 예정입니다.

Original Abstract

Reinforcement learning for flow-matching text-to-speech is complicated by deterministic ODE sampling: trajectory-level policy-gradient methods typically convert the ODE into an SDE and track per-step likelihood ratios, introducing stochastic perturbations and substantial overhead. We propose GROW, a group-relative advantage-weighted on-policy RL method that acts directly on the standard flow-matching objective. For each prompt, GROW samples a group of on-policy utterances, separately standardizes intelligibility and speaker-similarity rewards within the group, and combines them to reweight flow-matching regression. A Wasserstein-2 velocity penalty anchors the updated model to a frozen pretrained reference. A group-mean reward baseline is introduced to convert reward weighting into advantage weighting. For strong pretrained TTS models with concentrated rewards, positive exponential weighting is dominated by reward-agnostic self-imitation, whereas a zero-mean signed advantage preserves effective within-group credit assignment. Instantiated on DiTAR and evaluated on LibriSpeech and Seed-TTS EN/ZH, GROW reduces average WER from 2.016 to 1.558 and raises speaker similarity from 0.676 to 0.715 while keeping UTMOS. With 10-NFE training rollouts and 32-NFE evaluation, GROW retains comparable performance while training 2.9x faster than 32-NFE DiTAR-GRPO. We will open-source complete GROW codes, faithful DiTAR reproduction, and all model checkpoints.

0 Citations
0 Influential
5 Altmetric
25.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!