2605.29430v1 May 28, 2026 cs.AI

에이전트 기반 수정 및 의미 평가를 통한 인간과 유사한 상호 작용 음성 인식 연구

Towards Human-Like Interactive Speech Recognition With Agentic Correction and Semantic Evaluation

Kai Yu
Kai Yu
Citations: 542
h-index: 13
Zixu Jiang
Zixu Jiang
Citations: 208
h-index: 4
Wupeng Wang
Wupeng Wang
Citations: 156
h-index: 7
Xiangang Li
Xiangang Li
Citations: 61
h-index: 4
Xie Chen
Xie Chen
Citations: 250
h-index: 7
Yanqiao Zhu
Yanqiao Zhu
Citations: 181
h-index: 3
Zhifu Gao
Zhifu Gao
Citations: 2,583
h-index: 19
Peng Wang
Peng Wang
Citations: 0
h-index: 0
Qinyu Chen
Qinyu Chen
Citations: 47
h-index: 3
Xinjian Zhao
Xinjian Zhao
Citations: 4
h-index: 1
Xipeng Qiu
Xipeng Qiu
Citations: 20
h-index: 2

자동 음성 인식(ASR)은 인간-컴퓨터 상호작용의 핵심 구성 요소이며, LLM 기반 어시스턴트 및 에이전트의 중요한 전방 시스템입니다. 그러나 현재 대부분의 ASR 시스템은 여전히 단일 패스 방식으로 작동하며, 이는 오해가 반복적인 명확화와 수정을 통해 해결되는 인간 커뮤니케이션 방식과는 맞지 않습니다. 이러한 불일치는 발생한 의미에 중요한 오류를 수정하기 어렵게 만듭니다. 또한 WER 또는 CER과 같은 토큰 수준의 지표는 이러한 문제를 충분히 반영하지 못합니다. 이러한 제한 사항을 해결하기 위해, 우리는 extit{Interactive ASR}을 다중 턴(turn) 정제 작업으로 정의하고, 단일 패스 ASR 전방 시스템에 의미 수정, 의도 라우팅 및 추론 기반 편집 기능을 결합한 폐루프 프레임워크인 extbf{Agentic ASR}을 제안합니다. 또한 LLM 기반의 의미 평가 지표인 extbf{Sentence-level Semantic Error Rate}($S^2ER$)과 확장 가능하고 재현 가능한 벤치마킹을 위한 extbf{Interactive Simulation System}을 소개합니다. 다국어, 개체명 중심 및 코드 스위칭 벤치마크에서의 실험 결과, 반복적인 상호 작용은 의미 오류를 지속적으로 줄이며, 기존의 토큰 수준 지표보다 $S^2ER$ 측면에서 훨씬 더 큰 개선 효과를 보였습니다. 인간-AI 정렬 연구와 분석(ablation study)을 통해 의미 평가기의 신뢰성과 제안된 프레임워크의 견고성을 검증했습니다. 코드: https://interactiveasr.github.io/, 데모: https://i-asr.sjtuxlance.com/ 에서 확인할 수 있습니다.

Original Abstract

Automatic speech recognition (ASR) is a core component of human--computer interaction and an increasingly important front-end for LLM-based assistants and agents. However, most current ASR systems still follow a single-pass paradigm, which is poorly aligned with human communication, where misunderstandings are resolved through iterative clarification and refinement. This mismatch makes it difficult to correct meaning-critical errors once they occur. Meanwhile, token-level metrics such as WER or CER cannot adequately reflect such a problem. To address these limitations, we formulate \emph{Interactive ASR} as a multi-turn refinement task and propose \textbf{Agentic ASR}, a closed-loop framework that combines a single-pass ASR front-end with semantic correction, intent routing, and reasoning-based editing. We further introduce the \textbf{Sentence-level Semantic Error Rate} ($S^2ER$), an LLM-based semantic evaluation metric, together with an \textbf{Interactive Simulation System} for scalable and reproducible benchmarking. Experiments on multilingual, named-entity-intensive, and code-switching benchmarks show that iterative interaction consistently reduces semantic errors, with much larger gains in $S^2ER$ than in conventional token-level metrics. Human--AI alignment and ablation studies further validate the reliability of the semantic judge and the robustness of the proposed framework. The code is available at: https://interactiveasr.github.io/ and the live demo is available at https://i-asr.sjtuxlance.com/

0 Citations
0 Influential
9.5 Altmetric
47.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!