MOS(주관적 품질 평가) 지침을 활용한 발음 장애 심각도 평가 개선 연구
Augmenting Dysarthric Speech Severity Assessment with MOS Supervision
발음 장애는 언어의 명확성이 떨어지고 의사소통 효과가 감소하는 언어 장애입니다. 자동화된 발음 단편 수준의 발음 장애 평가는 대규모 음성 모니터링 및 치료 관련 분석을 지원할 수 있습니다. 그러나 이러한 시스템을 훈련하는 데에는 임상적으로 주석이 달린 발음 장애 음성의 부족으로 인해 어려움이 있습니다. 본 연구에서는 QualiSpeech 데이터베이스에서 얻은, 인간에 의해 평가된 평균 의견 점수(MOS) 레이블이 포함된 음성 합성 평가 데이터를 활용하여 발음 장애 평가를 개선하는 방법을 제안합니다. 실험 결과, 음성 합성 평가 데이터로 미세 조정하면 명확성과 자연스러움 예측 모두에서 성능이 꾸준히 향상되는 것으로 나타났습니다. 반면, 공동 훈련은 주로 자연스러움 측면에서 이점을 제공했습니다. 이러한 결과는 음성 합성의 인공적인 특징과 발음 장애 음성이 인식적으로 유사한 점을 시사하며, 음성 합성 평가 데이터베이스가 실용적인 보완 자료로서 활용되어 부족한 임상 주석에 대한 의존도를 줄일 수 있음을 보여줍니다.
Dysarthria is a speech disorder marked by reduced intelligibility and communicative effectiveness. Automatic utterance-level assessment of dysarthric speech can support scalable speech monitoring and therapy-related analysis. Yet training such systems is bottlenecked by the scarcity of clinically annotated dysarthric speech. This work proposes to augment dysarthric speech assessment using data from speech synthesis evaluations, specifically human-annotated utterances with Mean Opinion Score (MOS) labels from the QualiSpeech corpus. Experiments show that fine-tuning on speech synthesis assessment data consistently improves performance on both intelligibility and naturalness prediction, while joint training yields gains primarily on naturalness. These results suggest that synthesis artifacts and dysarthric speech share perceptual commonalities, and speech synthesis evaluation corpora offer a practical augmentation source that reduces reliance on scarce clinical annotations.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.