BareWave: 웨이브폼 기반 플로우 매칭 텍스트 음성 변환
BareWave: Waveform-Native Flow-Matching Text-to-Speech
중간 표현을 제거하고 별도로 학습된 디코딩 단계를 분리하는 것은 생성 모델링의 중요한 방향으로 자리 잡았습니다. 그러나 텍스트 음성 변환에서는 여전히 고품질 시스템이 종종 중간 음향 표현을 거쳐 웨이브폼을 합성하는 방식으로 구축됩니다. 본 연구에서는 플로우 매칭 텍스트 음성 변환에서 직접적인 텍스트-웨이브폼 생성을 위한 완전한 웨이브폼 기반 프레임워크인 BareWave를 제시합니다. 우리는 이 설정을 고려할 때 세 가지 학습상의 어려움이 있다고 판단했습니다. 첫째, 원본 웨이브폼 모델링은 강력한 사전 학습된 표현 체계를 갖추기 어렵습니다. 둘째, 학습의 각 단계는 서로 다른 노이즈 스케줄에서 효과를 얻을 수 있습니다. 셋째, 데이터 공간에서의 인지적 목표는 속도 공간 플로우 목표의 시간 구조를 자동으로 공유하지 않습니다. 그 결과, 직접적인 웨이브폼 학습은 효율적으로 최적화하기 어렵고, 일관된 방식으로 강력한 최종 성능을 달성하기 어렵고, 효과적인 인지적 개선을 통합하기 어렵습니다. 이러한 관점에 따라, 우리는 학습 시간 표현 정렬, 단계별 노이즈 스케줄링 및 속도 인식 인지적 정렬(VAPA)을 결합하는 직접적인 텍스트-웨이브폼 학습 프레임워크를 개발했습니다. 이 프레임워크는 사전 학습된 구성 요소 없이 단일 웨이브폼 기반 추론 경로를 유지합니다. 제로샷 음성 복제 실험 결과, 완전한 웨이브폼 기반 추론 경로에서 뛰어난 가독성, 화자 유사성 및 자연스러움을 달성할 수 있음을 보여주며, 이는 웨이브폼 기반 플로우 매칭 텍스트 음성 변환이 실용적인 접근 방식임을 뒷받침합니다. 오디오 데모가 포함된 프로젝트 페이지는 https://barewave.github.io/ 에서 확인할 수 있습니다.
Removing intermediate representations and separately trained decoding stages has become an important direction in generative modeling. In text-to-speech, however, high-quality systems are still commonly built through an intermediate acoustic representation before waveform synthesis. In this work, we present BareWave, a fully waveform-native framework for direct text-to-wave generation in flow-matching TTS. We consider this setting to raise three training challenges: raw-waveform modeling lacks a strong pretrained representational scaffold, different stages of training benefit from different noise schedules, and data-space perceptual objectives do not automatically share the temporal structure of the velocity-space flow objective. As a result, direct waveform training is hard to optimize efficiently, hard to push toward a strong final operating point with a fixed recipe, and hard to integrate effective perceptual refinement. Guided by this view, we develop a direct text-to-wave training framework that combines training-time representation alignment, staged noise scheduling, and velocity-aware perceptual alignment (VAPA), while preserving a single waveform-native inference path without pretrained components at test time. Experiments on zero-shot voice cloning show that strong intelligibility, speaker similarity, and naturalness can be achieved under a fully waveform-native inference path, supporting waveform-native flow-matching TTS as a practical direction. Project page with audio demos is available at https://barewave.github.io/.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.