2508.02038v5 Aug 04, 2025 cs.CL

마르코 보이스 기술 보고서

Marco-Voice Technical Report

Longyue Wang
Longyue Wang
Citations: 210
h-index: 6
Weihua Luo
Weihua Luo
Citations: 901
h-index: 14
Kaifu Zhang
Kaifu Zhang
Citations: 823
h-index: 13
Zhao Xu
Zhao Xu
Citations: 466
h-index: 10
Chenyang Lyu
Chenyang Lyu
Citations: 565
h-index: 13
Fengping Tian
Fengping Tian
Citations: 4
h-index: 1
Xuanfan Ni
Xuanfan Ni
Citations: 87
h-index: 5
Haoqin Sun
Haoqin Sun
Citations: 262
h-index: 10
Qingjuan Li
Qingjuan Li
Citations: 4
h-index: 1
Zhiqiang Qian
Zhiqiang Qian
Citations: 7
h-index: 2
Haijun Li
Haijun Li
Citations: 27
h-index: 3

본 논문은 음성 복제와 감정 제어 음성 합성 기능을 통합한 다기능 음성 합성 시스템을 제시합니다. 본 연구의 목표는 다양한 언어적, 감정적 맥락에서 화자 정보를 충실하게 유지하면서도 높은 수준의 표현력과 제어 가능성을 갖춘 자연스러운 음성 생성을 달성하는 데 있어 기존의 어려움을 해결하는 것입니다. 저희는 in-batch 대비 학습을 활용한 효과적인 화자-감정 분리 메커니즘과 부드러운 감정 제어를 위한 회전 감정 임베딩 통합 방법을 도입했습니다. 포괄적인 훈련 및 평가를 지원하기 위해, 저희는 6명의 전문 성우가 다양한 7가지 감정으로 표현한 총 10시간 분량의 만다린어 음성 데이터셋인 CSEMOTIONS을 구축했습니다. 광범위한 실험 결과, 저희 시스템인 Marco-Voice가 객관적 및 주관적 지표 모두에서 상당한 성능 향상을 달성하는 것을 보여줍니다. 종합적인 평가와 분석 결과, Marco-Voice는 음성 명료도와 감정 풍부도 측면에서 경쟁력 있는 성능을 제공하며, 이는 표현 신경망 음성 합성 분야의 중요한 진전을 의미합니다. 저희의 코드 및 데이터셋은 각각 https://github.com/AIDC-AI/Marco-Voice 와 https://huggingface.co/datasets/AIDC-AI/CSEMOTIONS 에서 공개적으로 이용할 수 있습니다.

Original Abstract

This paper presents a multifunctional speech synthesis system that integrates voice cloning and emotion control speech synthesis within a unified framework. The goal of this work is to address longstanding challenges in achieving highly expressive, controllable, and natural speech generation that faithfully preserves speaker identity across diverse linguistic and emotional contexts. Our approach introduces an effective speaker-emotion disentanglement mechanism with in-batch contrastive learning, enabling independent manipulation of speaker identity and eemotional style, as well as rotational emotional embedding integration method for smooth emotion control. To support comprehensive training and evaluation, we construct CSEMOTIONS, a high-quality emotional speech dataset containing 10 hours of Mandarin speech from six professional speakers across seven emotional categories. Extensive experiments demonstrate that our system, Marco-Voice, achieves substantial improvements in both objective and subjective metrics. Comprehensive evaluations and analysis were conducted, results show that MarcoVoice delivers competitive performance in terms of speech clarity and emotional richness, representing a substantial advance in the field of expressive neural speech synthesis. Our code and dataset are publicly available at https://github.com/AIDC-AI/Marco-Voice and https://huggingface.co/datasets/AIDC-AI/CSEMOTIONS respectively.

4 Citations
0 Influential
0 Altmetric
16.1 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!