2603.05231v1 Mar 05, 2026 cs.SD

오디오-텍스트 의미론적 보상을 활용한 테스트 시간 강화 학습을 통한 음성 인식 시스템의 강건성 향상

Boosting ASR Robustness via Test-Time Reinforcement Learning with Audio-Text Semantic Rewards

Tianxin Xie
Tianxin Xie
Citations: 90
h-index: 5
Li Liu
Li Liu
Citations: 11
h-index: 2
Linghan Fang
Linghan Fang
Citations: 7
h-index: 1

최근 음성 인식(ASR) 시스템(예: Whisper)은 상당한 정확도 향상을 이루었지만, 여전히 실제 환경에서 접하지 못한 데이터(큰 분포 변화를 가진 데이터), 특히 소음 환경 및 다양한 억양에 민감하게 반응합니다. 이러한 문제를 해결하기 위해, 테스트 시간 적응(TTA)은 ground truth 레이블 없이 추론 시 모델의 적응성을 향상시키는 데 큰 잠재력을 보여주며, 기존의 TTA 방법은 종종 pseudo-labeling 또는 엔트로피 최소화를 사용합니다. 그러나 이러한 방법들은 모델의 확신도를 학습 신호로 활용하기 때문에, 높은 확신도를 가진 오류를 강화하여 확인 편향을 초래할 수 있으며, 이는 적응을 저해할 수 있습니다. 이러한 한계점을 극복하기 위해, 우리는 인과적 개입에서 영감을 받은 새로운 테스트 시간 강화 적응 프레임워크인 ASR-TRA를 제시합니다. 구체적으로, 우리의 방법은 학습 가능한 디코더 프롬프트를 도입하고, 온도 조절된 확률적 디코딩을 사용하여 다양한 전사 후보를 생성합니다. 이러한 후보들은 오디오-텍스트 의미론적 정렬을 측정하는 보상 모델에 의해 점수화되며, 결과적인 피드백은 강화 학습을 통해 모델 및 프롬프트 파라미터를 업데이트하는 데 사용됩니다. 합성 노이즈가 포함된 LibriSpeech 데이터셋 및 L2 Arctic 억양 영어 데이터셋에 대한 종합적인 실험 결과, 우리의 방법이 기존의 TTA 기반 시스템보다 높은 정확도를 달성하면서도 낮은 지연 시간을 유지한다는 것을 보여줍니다. 추가적인 분석을 통해 오디오 및 언어 기반 보상을 결합하는 것이 효과적임을 확인했으며, 이는 우리의 방법이 향상된 안정성과 해석 가능성을 제공함을 강조합니다. 전반적으로, 우리의 접근 방식은 어려운 실제 환경에서 음성 인식 시스템을 배포하기 위한 실용적이고 강건한 솔루션을 제공합니다.

Original Abstract

Recently, Automatic Speech Recognition (ASR) systems (e.g., Whisper) have achieved remarkable accuracy improvements but remain highly sensitive to real-world unseen data (data with large distribution shifts), including noisy environments and diverse accents. To address this issue, test-time adaptation (TTA) has shown great potential in improving the model adaptability at inference time without ground-truth labels, and existing TTA methods often rely on pseudo-labeling or entropy minimization. However, by treating model confidence as a learning signal, these methods may reinforce high-confidence errors, leading to confirmation bias that undermines adaptation. To overcome these limitations, we present ASR-TRA, a novel Test-time Reinforcement Adaptation framework inspired by causal intervention. More precisely, our method introduces a learnable decoder prompt and utilizes temperature-controlled stochastic decoding to generate diverse transcription candidates. These are scored by a reward model that measures audio-text semantic alignment, and the resulting feedback is used to update both model and prompt parameters via reinforcement learning. Comprehensive experiments on LibriSpeech with synthetic noise and L2 Arctic accented English datasets demonstrate that our method achieves higher accuracy while maintaining lower latency than existing TTA baselines. Ablation studies further confirm the effectiveness of combining audio and language-based rewards, highlighting our method's enhanced stability and interpretability. Overall, our approach provides a practical and robust solution for deploying ASR systems in challenging real-world conditions.

1 Citations
0 Influential
2.5 Altmetric
13.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!