교차 데이터셋, 연령 및 성별 일반화: 저자원 아동 음성 인식 시스템의 미세 조정 전략에 대한 종합적인 분석
Cross-Dataset, Age, and Gender Generalization: A Comprehensive Analysis of Fine-Tuning Strategies for Low-Resource Children's ASR
운동 장애로 인한 발음의 부정확성이 야기하는 심각한 음향적 변이는 음성 인식의 주요 난제로 작용합니다. 기존 연구에서는 하이브리드 DNN/HMM 시퀀스 판별 학습을 통해 인식률 향상이 입증되었습니다. 본 논문은 다양한 음향 특징 조합을 분석하여, 각 음향 모델에 적합한 특징 선택 방법을 제시합니다. 특히 피치(음높이) 특징을 포함함으로써, 운동 장애를 가진 사람의 음성을 이용한 문장 인식 성능이 크게 향상되는 것을 확인했습니다. TORGO 데이터베이스를 활용한 체계적인 실험을 통해, 최첨단 Factorized Time Delay Neural Network (F-TDNN) 모델의 운동 장애 음성 인식 성능을 향상시킬 수 있는 가능성을 입증했습니다. 본 연구에서 제안하는 방법은 F-TDNN 모델을 사용하여 구현되었으며, 이전 연구에 비해 운동 장애된 사람의 단어 인식률이 4.65% 향상되고 문장 인식률이 4.63% 향상되었습니다. 이러한 성능 향상은 음성 변동성을 효과적으로 보완하며, 이는 연속적인 학습 예제 조각 간에 중첩되는 프레임 수를 신중하게 선택함으로써 달성되었습니다.
The challenge associated with recognizing dysarthric speech primarily arises from pronounced acoustic variability attributed to impaired articulatory precision. Past research has demonstrated improved recognition through the use of hybrid DNN/HMM sequence discriminative training. This paper presents a comprehensive investigation of various combinations of acoustic features tailored to different Acoustic Models, offering suitable feature selections for each. The incorporation of Pitch features notably improved recognition performance, especially for sentence recognition tasks involving dysarthric speech. Through a systematic examination of the TORGO database, we have demonstrated the potential to enhance the performance of the state-of-the-art Factorized Time Delay Neural Network (F-TDNN) model for recognizing dysarthric speech. Our methods, implemented with the F-TDNN model, resulted in a 4.65\% relative improvement in isolated word recognition and a 4.63\% relative improvement in sentence recognition for dysarthric speech, compared to previous research. This improvement effectively compensates for speech variability, attributable to our deliberate selection of the number of overlapping frames between consecutive training example chunks.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.