2606.24320v1 Jun 23, 2026 cs.SD

ZONOS2 기술 보고서

ZONOS2 Technical Report

Beren Millidge
Beren Millidge
Citations: 2,358
h-index: 26
Gabrielle E. Clark
Gabrielle E. Clark
Citations: 33
h-index: 3
Sofian Mejjoute
Sofian Mejjoute
Citations: 0
h-index: 0
Mohamed Osman
Mohamed Osman
Citations: 31
h-index: 3
George Close
George Close
Citations: 106
h-index: 6

본 논문에서는 ZONOS2 8B를 소개합니다. 이는 저희가 개발한 최신 TTS 모델로, 자연스러움, 운율, 그리고 음성 복제 정확도 측면에서 최고 수준의 성능을 달성했습니다. Zonos-v0.1에 비해 크기, 데이터, 그리고 학습 방법론에서 개선되었습니다. 새로운 Mixture-of-Experts (MoE) 구조를 사용하여 모델 크기를 1.6B에서 8B (실질적 파라미터 9억 개)로 확장하여 추론 지연 시간과 처리량을 향상시켰습니다. 또한, 새로운 데이터 처리 파이프라인을 통해 학습 데이터를 20만 시간 분량에서 6백만 시간 이상으로 늘리고, 자연스러움 및 음성 복제 정확도를 개선하기 위해 후처리 및 조건부 학습 방법을 간소화했습니다. ZONOS2 8B는 품질, 화자 유사도, 단어 오류율 (WER), 그리고 저희가 새로 개발한 TTS 벤치마크인 ZTTS1-Eval에서 평가되었으며, 최첨단 시스템과 경쟁력 있는 성능을 보이면서도 우수한 스트리밍 지연 시간을 유지합니다. 모델 가중치 및 예제 추론 코드를 Apache 2.0 라이선스에 따라 GitHub 및 Hugging Face에서 공개합니다.

Original Abstract

We present ZONOS2 8B, our latest TTS model, which achieves state-of-the-art naturalness, prosody, and voice cloning fidelity. We improve upon Zonos-v0.1 across scale, data, and training recipe. We scale the model from 1.6B to 8B total parameters (900M active) with a novel mixture-of-experts (MoE) backbone, improving inference latency and throughput. We expand our training corpus from 200K to over 6M hours using a new data processing pipeline, and we simplify our post-training and conditioning recipes to improve naturalness and voice cloning fidelity. We evaluate ZONOS2 8B on quality, speaker similarity, WER, and ZTTS1-Eval, our novel TTS benchmark, where it performs competitively with state-of-the-art systems while maintaining good streaming latency. We release our model weights and example inference code under an Apache 2.0 license on GitHub and Hugging Face.

0 Citations
0 Influential
13 Altmetric
65.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!