Sarashina2.2-TTS: 데이터 스케일링 및 타겟 데이터 증강을 통한 일본어 음성 생성 시 한자 다의음 문제 해결
Sarashina2.2-TTS: Tackling Kanji Polyphony in Japanese Speech Generation via Data Scaling and Targeted Data Synthesis
대규모 언어 모델(LLM) 기반 텍스트 음성 변환(TTS) 시스템은 고품질 음성 합성을 달성했지만, 대부분의 기존 시스템은 영어 및 중국어를 중심으로 개발되었습니다. 반면, 일본어는 상대적으로 연구가 부족하며, 맥락에 따라 의미가 달라지는 한자 다의음과 같은 고유한 언어적 난제가 충분히 해결되지 않았습니다. 본 논문에서는 이러한 문제점을 해결하기 위해 데이터 전략과 평가 방법론을 결합한 일본어 중심 LLM-TTS 시스템인 Sarashina2.2-TTS (https://github.com/sbintuitions/sarashina2.2-tts)를 소개합니다. 첫째, 36만 시간 분량의 음성 데이터를 활용하여 일본어 및 영어 데이터를 균형 있게 포함한 학습을 진행했습니다. 둘째, 일본 문화청에서 지정한 2,136개의 Joyo(일반 사용) 한자에 대한 모든 읽기 방식을 효율적으로 처리하기 위해 타겟 데이터 증강 파이프라인을 설계했습니다. 또한, Joyo Kanji Yomi Benchmark (https://github.com/sbintuitions/JoyoKanji-Yomi-Benchmark)를 소개하며, 이 벤치마크는 2,136개의 Joyo 한자와 그에 따른 4,378개의 읽기 방식을 포함합니다. 제시된 벤치마크와 함께, 생성된 음성을 기준 읽기와 비교하여 직교화 변동성을 제거하고 발음 정확도를 직접 측정하는 지표인 Kana-CER를 제안합니다. 실험 결과, 타겟 데이터 증강은 읽기 정확도를 크게 향상시키는 것으로 나타났습니다. 전체적으로 Sarashina2.2-TTS는 한자 수준의 읽기 정확도에서 최고 성능을 달성했으며, 일반적인 문장 수준의 발음에서는 최상의 기준 모델과 유사한 결과를 보였습니다. 또한, 제로샷 일본어 음성 합성에서 가장 높은 화자 유사도를 제공합니다. 더욱이, 교차 언어 평가 결과, Sarashina2.2-TTS는 프롬프트 언어와 관계없이 안정적인 일본어 발음을 유지하는 유일한 시스템임이 확인되었으며, 이는 균형 잡힌 학습 방식이 교차 언어적 견고성을 향상시키는 데 기여함을 보여줍니다.
While large language model (LLM)-based text-to-speech (TTS) systems have achieved high-quality speech synthesis, most existing systems focus on English and Chinese. Japanese, however, remains under-explored, and its unique linguistic challenges, such as widespread context-dependent kanji polyphony, have yet to be adequately tackled. Here we introduce Sarashina2.2-TTS (https://github.com/sbintuitions/sarashina2.2-tts), a Japanese-centric LLM-TTS system that tackles these challenges through a dual approach: data strategy and evaluation methodology. First, we scale training to approximately 361k hours of speech, incorporating a balanced mix of Japanese and English data. Furthermore, we design a targeted data augmentation pipeline covering all 2,136 Joyo (regular-use) kanji designated by Japan's Agency for Cultural Affairs to efficiently address kanji polyphony disambiguation. Second, we introduce the Joyo Kanji Yomi Benchmark (https://github.com/sbintuitions/JoyoKanji-Yomi-Benchmark), covering all 2,136 Joyo kanji and their 4,378 readings. Alongside this benchmark, we propose Kana-CER, a metric that compares synthesized speech against reference readings in the kana space, eliminating orthographic variations to directly measure pronunciation correctness. Experiments demonstrate that our targeted data augmentation significantly improves reading accuracy. Overall, Sarashina2.2-TTS achieves state-of-the-art kanji-level reading accuracy and matches top baselines on general sentence-level pronunciation, while delivering the highest speaker similarity in zero-shot Japanese speech synthesis. Furthermore, cross-lingual evaluation reveals that Sarashina2.2-TTS is the only system that maintains stable Japanese pronunciation regardless of the prompt language, confirming that our balanced training approach improves cross-lingual robustness.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.