2606.13464v1 Jun 11, 2026 cs.CL

온톨로지 메모리 기반 음성 인식 교정: 긴 텍스트-음성 혼합 대화

Ontology Memory-Augmented ASR Correction for Long Text-Speech Interleaved Conversations

Meishan Zhang
Meishan Zhang
Citations: 790
h-index: 12
Zulong Chen
Zulong Chen
Citations: 2
h-index: 1
Yunxin Li
Yunxin Li
Harbin Institute of Technology, Shenzhen
Citations: 1,113
h-index: 13
Xinxin Li
Xinxin Li
Citations: 1
h-index: 1
Huiyao Chen
Huiyao Chen
Citations: 70
h-index: 5
Zhibo Ren
Zhibo Ren
Citations: 4
h-index: 1
Xiaoqing Hu
Xiaoqing Hu
Citations: 4
h-index: 1
Min Zhang
Min Zhang
Citations: 104
h-index: 6

음성 인식(ASR) 교정은 전통적으로 독립적인 발화 또는 짧은 지역적 맥락에 초점을 맞춰 왔습니다. 그러나 텍스트와 음성이 장시간의 상호작용에서 점점 더 복잡하게 결합됨에 따라, ASR 교정은 대화 수준의 문맥 정보가 필요합니다. 기존 ASR 교정 방법은 종종 현재 가설에 의존하거나 원시 대화 기록을 연결하는 방식을 사용합니다. 이러한 맥락에서는 희소한 교정 정보를 과도한 정보와 노이즈 속에서 찾기가 어려울 수 있습니다. 이러한 문제점을 해결하기 위해, 우리는 긴 텍스트-음성 혼합 대화를 위한 온톨로지 메모리 기반 ASR 교정 프레임워크를 제안합니다. 이 프레임워크는 이전 상호작용 기록을 동적으로 업데이트 가능한 온톨로지 메모리에 구성하며, 엔티티, 전문 용어, 표면 변형, 잠재적인 ASR 오인, 그리고 의미 관계 등을 검색 가능한 노드로 저장하여 문맥 기반의 교정을 지원합니다. 이 설정을 평가하기 위해, 우리는 장거리 ASR 교정을 위한 문맥 기반 정보를 제공하는 MAGIC-RAMC 데이터 세트를 기반으로 RAMC-Corr 데이터 세트를 구축했습니다. RAMC-Corr에 대한 실험 결과, 제안된 방법은 10개의 주요 설정 조합 중 9개에서 직접 교정보다 성능이 우수했으며, 문맥 의존적인 ASR 오류에 대해 더욱 선택적이고 근거 기반의 교정을 수행하는 것으로 나타났습니다.

Original Abstract

Automatic speech recognition (ASR) correction has traditionally focused on isolated utterances or short local contexts. However, as text and speech become increasingly interleaved in long interactions, ASR correction requires conversation-level contextual evidence. Existing ASR correction methods often rely on the current hypothesis or concatenate raw dialogue history. In such contexts, sparse correction evidence can be difficult to locate amid redundancy and noise. Addressing these challenges, we propose an ontology memory-augmented ASR correction framework for long text-speech interleaved conversations. The framework organizes preceding interaction history into a dynamically updatable ontology memory, where entities, terminology, surface variants, potential ASR confusions, and semantic relations are stored as retrievable nodes for context-grounded correction. To evaluate this setting, we construct RAMC-Corr, a dataset derived from MAGIC-RAMC for long-range ASR correction with grounded context. Experiments on RAMC-Corr show that our method improves over direct correction in 9 out of 10 paired backbone-setting combinations and encourages more selective and evidence-grounded corrections for context-dependent ASR errors.

0 Citations
0 Influential
6.5 Altmetric
32.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!