2607.28175v1 Jul 30, 2026 cs.AI

AgenticASR: 에이전트 기반 접근 방식을 통한 실제 환경에서의 음성 인식 성능 향상

AgenticASR: Refining Speech Recognition in Real-World Scenarios via an Agentic Approach

Yanqiao Zhu
Yanqiao Zhu
Citations: 181
h-index: 3
Kai Yu
Kai Yu
Citations: 730
h-index: 9
Bin Qiang
Bin Qiang
Citations: 26
h-index: 3
Jiaying Chi
Jiaying Chi
Citations: 0
h-index: 0
Zixuan Jiang
Zixuan Jiang
Citations: 69
h-index: 2
Xie Chen
Xie Chen
Citations: 1,106
h-index: 13

음성 인식(ASR) 기술은 상당한 수준으로 정확도가 향상되었지만, 완벽하게 기록된 텍스트가 항상 사용하기 편리하지는 않습니다. ASR 결과물에는 불필요한 중복, 반복, 수정사항 등이 포함되어 있어 읽기 어려움을 야기하고, 화자의 의도를 명확하게 전달하지 못하며, 미해결 또는 폐기된 내용이 후속 작업에 영향을 줄 수 있습니다. 기존의 음성-텍스트 변환 방법은 이미 완료된 오디오나 텍스트를 처리하지만, 이후의 발화가 이전 내용을 어떻게 해석해야 하는지에 대한 변경 사항을 반영할 수 없습니다. 따라서 본 연구에서는 불필요한 요소들을 제거하고, 자기 수정 사항을 해결하며, 화자의 의도를 보존하면서 깨끗한 텍스트를 생성하는 '에이전트 기반 음성 인식(AgenticSR)'이라는 새로운 접근 방식을 제안합니다. AgenticASR은 ASR과 Refiner 아키텍처를 결합하여 오디오가 입력되는 동안 제한된 범위의 컨텍스트를 반복적으로 변환하고 해당 출력 부분을 수정하는 방식으로 작동하며, 이를 통해 임의의 길이의 스트림에서 지속적인 음성 생성 및 수정을 가능하게 합니다. 또한, 세밀한 평가 기준을 포함하는 이중 언어 벤치마크인 AASR-Bench를 새롭게 제안합니다. 다양한 ASR 시스템에 대한 실험 결과, AgenticASR은 평가된 시스템 중에서 가장 높은 AASR-Bench 점수를 기록했습니다. 인간과 AI의 합의 연구 결과는 평가 기준이 독립적인 전문가의 판단과 일치함을 보여줍니다. 추가적으로 Refiner의 성능, 컨텍스트 길이, 온라인 및 오프라인 추론 간의 품질-지연 균형에 대한 분석을 수행했습니다. 이러한 결과를 종합하면 AgenticASR은 지속적인 음성 인식 과정에서 화자의 의도를 보존하면서 깨끗한 텍스트를 생성하는 실용적인 프레임워크임을 알 수 있습니다. 코드, AASR-Bench 및 데모는 https://github.com/AnXMuy/AgenticASR 에서 확인할 수 있습니다.

Original Abstract

Automatic speech recognition (ASR) has achieved substantial gains in transcription accuracy, yet verbatim transcription does not necessarily produce readily usable text. It retains fillers, repetitions, false starts, and self-corrections that increase reading effort, obscure the speaker's final intent, and propagate unresolved or abandoned content to downstream tasks. Existing spoken-to-written methods process completed audio or transcripts but cannot revise emitted text when later speech changes how preceding content should be interpreted. We therefore formulate Agentic Speech Recognition (AgenticSR), an audio-to-clean-text task that removes disfluencies, resolves self-corrections, and normalizes written form while preserving the speaker's final intent. AgenticASR implements this task through an ASR--Refiner architecture that repeatedly transforms a bounded active context and replaces its corresponding output span as audio arrives. This enables continual emission and revision over streams of arbitrary duration. We also introduce AASR-Bench, a bilingual benchmark with fine-grained atomic rubrics. Across multiple ASR front ends, AgenticASR attains the highest AASR-Bench scores among evaluated systems. A human--AI agreement study shows that rubric-based judgments align with independent expert assessments. Ablations characterize Refiner capacity, context length, and the quality--latency trade-off between online and offline inference. Together, these results establish AgenticASR as a practical framework for intent-preserving clean transcription during ongoing speech. Code, AASR-Bench, and a demo will be released at https://github.com/AnXMuy/AgenticASR.

0 Citations
0 Influential
28.047189562171 Altmetric
0.0 Score
Original PDF
4

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!