SE-AGCNet: 회의 환경에서 음성 향상 및 음량 조절을 위한 통합 프레임워크
SE-AGCNet: An End-to-End Framework for Joint Speech Enhancement and Loudness Control in Meeting Scenarios
기존 오디오 파이프라인은 일반적으로 음성 향상(SE)과 자동 이득 제어(AGC)를 개별적인 모듈로 처리하여 전체 성능에 제한을 둡니다. 예를 들어, SE 전에 AGC를 적용하면 배경 소음을 의도치 않게 증폭시킬 수 있으며, SE를 우선적으로 수행하면 낮은 음량의 음성을 과도하게 감쇠시킬 수 있습니다. 이러한 한계점을 해결하기 위해, 우리는 음성 향상과 자동 이득 제어를 동시에 최적화하는 통합 프레임워크인 SE-AGCNet을 제안합니다. 특히 회의 환경에서 발생하는 상당한 음량 변화에 맞춰 설계된 SE-AGCNet은 두 작업 간의 시너지 효과를 활용합니다. 즉, SE는 낮은 음량의 음성을 보존하여 AGC 구성 요소가 효과적인 음량 조정을 수행할 수 있도록 지원합니다. 또한, 우리는 특수 데이터 생성 파이프라인인 SE-AGC-DataGen을 제안하고, 표준화된 음량 평가 지표인 통합 음량(LUFS), 단기 음량(St LUFS) 및 LRA를 포함했습니다. 실험 결과, SE-AGCNet은 경쟁적인 기준 모델보다 음성 품질과 ASR 정확도를 향상시키면서도 목표 음량을 지속적으로 달성하는 것으로 나타났습니다.
Conventional audio pipelines typically treat speech enhancement (SE) and automatic gain control (AGC) as discrete modules, which often limits overall performance. For instance, applying AGC before SE may inadvertently amplify background noise, while prioritizing SE tends to over-suppress low-volume speech. To address these limitations, we propose SE-AGCNet, an end-to-end framework that jointly optimizes SE and AGC. Tailored for meeting scenarios with significant volume variations, SE-AGCNet leverages the synergy between the two tasks: SE preserves quiet speech, thereby facilitating effective volume adjustment by the AGC component. Furthermore, we propose a specialized data simulation pipeline, SE-AGC-DataGen, and incorporate standardized loudness evaluation metrics: integrated loudness (LUFS), short-term loudness (St LUFS), and LRA. Experiments show that SE-AGCNet consistently achieves target loudness while improving speech quality and ASR accuracy over competitive baselines.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.