구음장애 음성 인식에 대한 체계적 연구: 스펙트럴 특징과 음향 모델
Systematic Study of Dysarthric Speech Recognition: Spectral Features and Acoustic Models
구음장애 음성 인식이 어려운 주된 이유는 발음 기관의 기능 저하로 인한 현저한 음향 변동성 때문입니다. 이전 연구에서는 하이브리드 DNN/HMM 시퀀스 차등 학습을 통해 인식률 향상이 입증되었습니다. 본 논문은 다양한 음향 특징과 음향 모델 조합에 대한 종합적인 조사를 제시하며, 각 모델에 적합한 특징 선택 방법을 제공합니다. 특히 음높이(Pitch) 특징을 포함시키면, 구음장애 음성을 이용한 문장 인식 작업에서 인식 성능이 크게 향상되었습니다. TORGO 데이터베이스를 활용한 체계적인 분석을 통해, 최첨단 Factorized Time Delay Neural Network (F-TDNN) 모델의 구음장애 음성 인식 성능을 향상시킬 수 있음을 입증했습니다. F-TDNN 모델을 기반으로 구현된 방법은, 이전 연구에 비해 구음장애 음성에 대한 단어 인식률을 4.65% 상대적으로 향상시키고, 문장 인식률을 4.63% 상대적으로 향상시켰습니다. 이러한 성능 향상은 음성 변동성을 효과적으로 보완하며, 이는 연속적인 학습 예제 조각 간에 중첩되는 프레임 수를 신중하게 선택함으로써 달성되었습니다.
The challenge associated with recognizing dysarthric speech primarily arises from pronounced acoustic variability attributed to impaired articulatory precision. Past research has demonstrated improved recognition through the use of hybrid DNN/HMM sequence discriminative training. This paper presents a comprehensive investigation of various combinations of acoustic features tailored to different Acoustic Models, offering suitable feature selections for each. The incorporation of Pitch features notably improved recognition performance, especially for sentence recognition tasks involving dysarthric speech. Through a systematic examination of the TORGO database, we have demonstrated the potential to enhance the performance of the state-of-the-art Factorized Time Delay Neural Network (F-TDNN) model for recognizing dysarthric speech. Our methods, implemented with the F-TDNN model, resulted in a 4.65\% relative improvement in isolated word recognition and a 4.63\% relative improvement in sentence recognition for dysarthric speech, compared to previous research. This improvement effectively compensates for speech variability, attributable to our deliberate selection of the number of overlapping frames between consecutive training example chunks.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.