2606.22811v1 Jun 22, 2026 cs.CL

Bagpiper-TTS: 자연어 기반의 범용 음성 합성 시스템

Bagpiper-TTS: Natural Language Guided Universal Speech Synthesis

Chao-Han Huck Yang
Chao-Han Huck Yang
Citations: 94
h-index: 3
Shinji Watanabe
Shinji Watanabe
Citations: 82
h-index: 6
Siddhant Arora
Siddhant Arora
Citations: 1,464
h-index: 18
Jinchuan Tian
Jinchuan Tian
Citations: 0
h-index: 0
Haoran Wang
Haoran Wang
Citations: 2
h-index: 1
Takashi Maekaku
Takashi Maekaku
Citations: 338
h-index: 8
Keita Goto
Keita Goto
Citations: 14
h-index: 2
Jin Sakuma
Jin Sakuma
Citations: 352
h-index: 6
Yusuke Shinohara
Yusuke Shinohara
Citations: 15
h-index: 2

기존의 TTS(Text-to-Speech) 시스템은 일반적으로 엄격한 입력 형식과 미리 정의된 메타데이터 슬롯에 의존하여, 사용자의 다양한 요구 사항을 충족하는 데 한계가 있습니다. 본 논문에서는 자연어 사용자 요청을 처리할 수 있는 범용 음성 합성 시스템인 Bagpiper-TTS를 소개합니다. Bagpiper-TTS는 주어진 자연어 프롬프트를 기반으로 사용자의 의도를 파악하여 풍부한 캡션을 생성합니다. 이 캡션은 음성 합성을 안내하며, 여기에는 정확한 발음과 함께 미묘한 메타데이터 정보가 포함됩니다. 또한, 저희 모델은 기존 TTS 애플리케이션 외에도 다중 화자, 의도-음성 변환, 역할극 합성, 노래 목소리 합성 등 광범위한 작업을 지원합니다. 실험 결과는 Bagpiper-TTS가 Seed-TTS-Eval 벤치마크에서 1.7%의 단어 오류율(WER)을 달성했으며, LLM 기반 평가 및 인간 주관적 평가를 통해 다양한 애플리케이션에서 전문 모델과 동등한 성능을 보임을 입증합니다.

Original Abstract

Classical TTS systems typically rely on rigid input formats and predefined metadata slots, limiting their ability to fulfill flexible user requirements. This paper introduces Bagpiper-TTS, a universal speech synthesis system that deals with diverse natural language user requests. Given a natural language prompt, Bagpiper-TTS first reasons over the users' intent to derive a rich caption, i.e., a comprehensive textual blueprint encompassing both transcription and nuanced metadata. Subsequently, this caption guides the synthesis of the target speech. Our model inherently supports a broad spectrum of tasks besides classical TTS applications, including multi-talker, intent-to-speech, role-play synthesis, singing voice synthesis, and more. Experimental results demonstrate that Bagpiper-TTS achieves an 1.7% Word Error Rate (WER) on the Seed-TTS-Eval benchmark and match the performance of dedicated models in both LLM-as-a-judge and human subjective evaluations across multiple applications.

0 Citations
0 Influential
9 Altmetric
45.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!