2606.19797v1 Jun 18, 2026 eess.AS

발음 장애 음성에 대한 종단 간 음성 인식 성능 향상을 위한 도메인 내 데이터 증강

Improving End-to-End Speech Recognition for Dysarthric Speech through In-Domain Data Augmentation

Sudarsana Reddy Kadiri
Sudarsana Reddy Kadiri
University of Southern California
Citations: 1,522
h-index: 22
Paban Sapkota
Paban Sapkota
Citations: 18
h-index: 3
H. Kathania
H. Kathania
Citations: 459
h-index: 13
Shrikanth S. Narayanan
Shrikanth S. Narayanan
Citations: 59
h-index: 4

발음 장애 음성 인식은 발음 장애를 가진 개인들의 효과적인 의사소통을 지원하는 데 매우 중요합니다. 그러나 발음 장애 정도의 차이가 크고 데이터 가용성이 제한적이기 때문에, 정확하게 발음 장애 음성을 인식하는 것은 상당한 어려움을 야기합니다. 본 연구에서는 엔드 투 엔드 사전 학습 모델인 Wav2Vec2를 미세 조정하여 발음 장애 자동 음성 인식(ASR) 시스템을 위한 데이터 증강 기법을 탐색하며, 특히 발음 장애 정도에 초점을 맞춥니다. 발음 장애 음성을 위한 사전 학습 ASR 시스템의 미세 조정에 필요한 풍부한 데이터 확보 문제와 데이터 부족 문제를 해결하기 위해, 발음 장애의 다양한 측면을 고려하여 Speaking-Rate Modification (SRM), Pitch Modification (PM), Formant Modification (FM) 및 vocal tract Length Perturbation (VTLP)의 네 가지 주요 데이터 증강 방법을 조사했습니다. 본 연구에서는 개별적으로 미세 조정된 Wav2Vec2 모델을 각 발음 장애 정도 클래스에 대해 기본 시스템으로 사용했습니다. 또한, 증강 데이터를 사용하여 ASR 모델의 발음 장애 정도에 따른 특화된 미세 조정을 수행했습니다. 실험 결과는 각 증강 기법이 발음 장애 정도에 따라 뚜렷한 효과를 나타냄을 보여줍니다. extit{낮은} 정도(9.02%)와 extit{중간} 정도(38.11%)의 경우 SRM ($s$=0.8)에서 가장 좋은 WER(Word Error Rate, 단어 오류율) 값을 얻었으며, extit{높은} 정도(55.15%)의 경우 PM ($τ$=0.8)에서 가장 좋은 결과를 보였습니다. 이는 각각 30.02%, 16.64% 및 15.47%의 상대적인 성능 향상을 나타냅니다. 이러한 결과는 제안된 증강 방법들이 발음 장애 ASR 성능을 개선하는 데 효과적임을 확인합니다.

Original Abstract

Dysarthric speech recognition is crucial for facilitating effective communication among individuals with dysarthria. However, accurately recognizing dysarthric speech poses significant challenges due to varying severity levels and limited data availability. In this paper, we explore data augmentation techniques for dysarthric automatic speech recognition (ASR) systems by fine-tuning the End-to-End pre-trained Wav2Vec2 model, with a specific focus on severity levels. To address the challenges of data scarcity and the need for extensive data in fine-tuning pre-trained ASR systems for dysarthric speech, we investigate four prominent data augmentation methods: Speaking-Rate Modification (SRM), Pitch Modification (PM), Formant Modification (FM), and vocal tract Length Perturbation (VTLP), tailored to different aspects of dysarthria. The study uses individually fine-tuned Wav2Vec2 models for each severity class as baseline systems. Additionally, we conducted severity-specific fine-tuning of the ASR model using augmented data. Results demonstrate distinct efficacy patterns for each augmentation technique across severity levels. The best WERs were achieved with SRM ($s$=0.8) for \textit{low} (9.02\%) and \textit{medium} (38.11\%) severities, and with PM ($τ$=0.8) for \textit{high} severity (55.15\%), reflecting relative improvements of 30.02\%, 16.64\%, and 15.47\%, respectively. These results confirm the effectiveness of the augmentation methods in improving dysarthric ASR performance.

5 Citations
0 Influential
11 Altmetric
60.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!