Bagpiper-TTS: 자연어 기반의 범용 음성 합성 시스템
Bagpiper-TTS: Natural Language Guided Universal Speech Synthesis
기존의 TTS(Text-to-Speech) 시스템은 일반적으로 엄격한 입력 형식과 미리 정의된 메타데이터 슬롯에 의존하여, 사용자의 다양한 요구 사항을 충족하는 데 한계가 있습니다. 본 논문에서는 자연어 사용자 요청을 처리할 수 있는 범용 음성 합성 시스템인 Bagpiper-TTS를 소개합니다. Bagpiper-TTS는 주어진 자연어 프롬프트를 기반으로 사용자의 의도를 파악하여 풍부한 캡션을 생성합니다. 이 캡션은 음성 합성을 안내하며, 여기에는 정확한 발음과 함께 미묘한 메타데이터 정보가 포함됩니다. 또한, 저희 모델은 기존 TTS 애플리케이션 외에도 다중 화자, 의도-음성 변환, 역할극 합성, 노래 목소리 합성 등 광범위한 작업을 지원합니다. 실험 결과는 Bagpiper-TTS가 Seed-TTS-Eval 벤치마크에서 1.7%의 단어 오류율(WER)을 달성했으며, LLM 기반 평가 및 인간 주관적 평가를 통해 다양한 애플리케이션에서 전문 모델과 동등한 성능을 보임을 입증합니다.
Classical TTS systems typically rely on rigid input formats and predefined metadata slots, limiting their ability to fulfill flexible user requirements. This paper introduces Bagpiper-TTS, a universal speech synthesis system that deals with diverse natural language user requests. Given a natural language prompt, Bagpiper-TTS first reasons over the users' intent to derive a rich caption, i.e., a comprehensive textual blueprint encompassing both transcription and nuanced metadata. Subsequently, this caption guides the synthesis of the target speech. Our model inherently supports a broad spectrum of tasks besides classical TTS applications, including multi-talker, intent-to-speech, role-play synthesis, singing voice synthesis, and more. Experimental results demonstrate that Bagpiper-TTS achieves an 1.7% Word Error Rate (WER) on the Seed-TTS-Eval benchmark and match the performance of dedicated models in both LLM-as-a-judge and human subjective evaluations across multiple applications.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.