2608.05759v1 Aug 06, 2026 cs.CL

새로운 단어 인식 방법 비교 연구: 문맥 편향 기법과 음성 거대 언어 모델

How to Recognize New Words: A Comparison Between Context Biasing Methods and Speech LLMs

Christian Huber
Christian Huber
Citations: 97
h-index: 6
Alexander Waibel
Alexander Waibel
Citations: 50
h-index: 3

자동 음성 인식(ASR)에서 새로운 단어와 희귀 단어 (예: 고유 명사, 약어, 특정 분야 전문 용어 등)를 인식하는 것은 여전히 중요한 과제입니다. 본 연구에서는 두 가지 전략을 비교합니다. 첫 번째는 문맥 편향 기법으로, ASR 모델을 확장하여 추론 과정에서 단어 목록을 제공할 수 있도록 하는 방법입니다. 두 번째는 음성 거대 언어 모델(LLM)에 직접적인 문맥 정보를 입력하는 방법입니다. Whisper를 기반으로 한 두 가지 문맥 편향 기법과 세 개의 음성 LLM을 사용하여 읽은 음성과 자연스러운 음성에 대한 성능을 평가하고, 편향된 단어 오류율(WER), 편향되지 않은 단어 오류율, 그리고 전체적인 단어 오류율을 보고합니다. 문맥 편향 기법은 편향된 WER을 최대 88%까지 감소시키는 효과를 보였으며, 다른 단어에 대한 영향은 미미했습니다. 음성 LLM은 읽은 음성에서는 뛰어난 성능을 보이지만, 자연스러운 음성에 대해서는 일반화 능력이 떨어지는 경향이 있으며, 주의 산만 요소의 수와 프롬프트 단어 순서에 민감하게 반응합니다. 본 연구는 이러한 결과를 바탕으로 방법 선택에 대한 지침을 제공하고자 합니다.

Original Abstract

Recognizing new and rare words - named entities, acronyms, domain specific special words, and other items scarce in training data - remains a key challenge for automatic speech recognition (ASR). We compare two strategies for this: context biasing methods, where an ASR model is extended such that during inference a word list can be supplied, and speech large language models (LLMs) prompted with context directly. We evaluate two context biasing methods based on Whisper against three speech LLMs across read and non-read speech, reporting biased, unbiased, and overall word error rate (WER). The context biasing methods cut biased WER by up to 88% relative while leaving other words largely unaffected. Speech LLMs excel on read speech but generalize less well to non-read speech, and prove sensitive to distractor count and prompt word order. We characterize the resulting trade-offs to guide method selection.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!