음성 번역에서의 불유창 요소의 역할
The Role of Disfluencies in Speech Translation
현재 음성 번역 시스템, 특히 SpeechLLM은 정제된 텍스트로 학습되며, 일반적으로 '어', '음'과 같은 채움 일시정지나 잘못 시작한 발화와 같은 불유창 요소를 번역하는 대신 제거하는 경향이 있습니다. 본 연구에서는 이러한 현상이 정보 손실을 초래한다는 것을 보여줍니다. 불유창 요소는 의미를 담고 있으며, 음성이 정제될 때 이 의미가 사라집니다. 이를 체계적으로 연구하기 위해, 영어 음성을 8개의 타겟 언어로 번역하고 불유창 요소를 주석으로 달아놓은 데이터셋인 Uh-Mazing을 개발했습니다. 여러 언어와 아키텍처를 대상으로 분석한 결과, 채움 일시정지나 담화 지표가 아닌 잘못 시작된 발화 및 자기 수정이 대부분의 번역 품질 저하를 유발하며, 불유창 요소를 보존하지 못하는 모델은 이를 오번역하기보다는 생략하는 경향을 보이는 것을 확인했습니다. 또한, 별도의 재학습 없이 추론 과정에서의 디코딩 방식을 통해 이러한 문제를 완화할 수 있으며, 개발된 데이터셋과 코드를 공개합니다.
Current speech translation systems, including SpeechLLMs, are trained on cleaned text and tend to strip disfluencies like filled pauses and false starts rather than translate them. We show this comes at a cost: disfluencies carry meaning that gets lost when speech is cleaned up. To study this systematically, we introduce Uh-Mazing, a benchmark of human-translated, disfluency-annotated Switchboard speech covering English into eight target languages. Across these languages and several architectures, we find that false starts and self-repairs, not filled pauses or discourse markers, drive most of the translation-quality loss, and that models which fail to preserve a disfluency tend to omit it rather than mistranslate it. We show inference-time decoding can mitigate this without retraining, and release the benchmark and code.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.