2606.05852v1 Jun 04, 2026 cs.SD

UniVoice: 음성 및 노래 목소리 생성 모델

UniVoice: A Unified Model for Speech and Singing Voice Generation

Junjie Zheng
Junjie Zheng
Citations: 40
h-index: 5
Huixin Xue
Huixin Xue
Citations: 41
h-index: 2
Chaofan Ding
Chaofan Ding
Citations: 67
h-index: 6
Shihong Ren
Shihong Ren
Citations: 44
h-index: 4
Haohe Liu
Haohe Liu
Citations: 299
h-index: 6
Zihao Chen
Zihao Chen
Citations: 55
h-index: 5

텍스트 음성 변환(TTS)과 노래 목소리 합성(SVS)은 모두 기호 입력을 기반으로 인간의 발성 오디오를 생성하는 것을 목표로 하지만, 생성 과정에 서로 다른 요구 사항을 적용합니다. 음성 생성은 유연하고 언어 기반의 운율에 의존하는 반면, 노래 생성은 명시적인 멜로디 제어와 정확한 리듬 정렬이 필요합니다. 이러한 차이점 때문에 자연스러운 음성과 제어 가능한 노래를 모두 생성할 수 있는 단일 모델을 학습하기 어렵습니다. 왜냐하면 멜로디 관련 조건은 노래에 강력하게 작용해야 하지만, 음성의 운율을 제한해서는 안 되기 때문입니다. 본 논문에서는 조건부 플로우 매칭 기반의 통합 음성 및 노래 목소리 생성 프레임워크인 UniVoice를 제시합니다. UniVoice는 단일하고 획일적인 조건 표현을 사용하는 대신, 콘텐츠, 멜로디, 그리고 음색으로 조건을 분해하여 모달리티에 적합한 인코더로 각각 인코딩하고 공유되는 Diffusion Transformer (DiT) 백본에서 사용합니다. 노래 생성의 경우, 멜로디 조건은 MIDI 노트 시퀀스로 표현되며, 음성의 경우에는 학습된 null 멜로디 토큰으로 대체되어 모델이 언어적 및 음향적 맥락으로부터 운율을 추론할 수 있도록 합니다. 이러한 설계는 노래에 대한 명시적인 멜로디 제어를 유지하면서 음성에 멜로디 제약 조건을 적용할 필요성을 없앱니다. 또한, 본 논문에서는 조건부 플로우에서의 멜로디 마진화 근사치를 나타내는 null 멜로디 토큰을 분석합니다. UniVoice는 30,000시간의 음성 데이터와 35,000시간의 노래 데이터로 학습되었으며, F5-TTS (5.21%) 및 CosyVoice3 (5.30%)과 같은 특수 TTS 시스템에 필적하는 5.26%의 음성 PER을 달성했습니다. 노래 생성 측면에서 UniVoice는 16.22%의 PER을 달성하여 통합 기준 모델인 Vevo1.5 (24.72%)보다 성능이 우수합니다.

Original Abstract

Text-to-speech (TTS) and singing voice synthesis (SVS) both aim to generate human vocal audio from symbolic inputs, but they impose different requirements on the generation process. Speech generation relies on flexible, language-driven prosody, whereas singing generation requires explicit melody control and accurate rhythmic alignment. This mismatch makes it challenging to train a single model that can generate both natural speech and controllable singing, since melody-related conditions should strongly constrain singing but should not restrict speech prosody. We present UniVoice, a unified speech and singing voice generation framework based on conditional flow matching. Instead of using a single undifferentiated conditioning representation, UniVoice factorizes the condition into content, melody, and timbre, which are encoded by modality-appropriate encoders and consumed by a shared Diffusion Transformer (DiT) backbone. For singing, the melody condition is represented by MIDI note sequences; for speech, it is replaced with a learned null melody token, allowing the model to infer prosody from linguistic and acoustic context. This design preserves explicit melody control for singing while avoiding the need to impose melody constraints on speech. We further analyze the null melody token as an approximation to melody marginalization in the conditional flow. Trained on 30k hours of speech and 35k hours of singing data, UniVoice achieves a speech PER of 5.26\%, comparable to dedicated TTS systems such as F5-TTS (5.21\%) and CosyVoice3 (5.30\%). On singing generation, UniVoice achieves a PER of 16.22\%, outperforming the unified baseline Vevo1.5 (24.72\%).

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!