2607.26410v1 Jul 29, 2026 cs.CL

에이전트 기반 음성 인식 시스템을 위한 보이스 메모리

Voice Memory for Agentic Speech Recognition

Zhehuai Chen
Zhehuai Chen
Citations: 594
h-index: 12
Boris Ginsburg
Boris Ginsburg
Citations: 1,674
h-index: 16
Zih-Ching Chen
Zih-Ching Chen
Citations: 9
h-index: 2
Chao-Han Huck Yang
Chao-Han Huck Yang
Citations: 108
h-index: 4
Piotr Żelasko
Piotr Żelasko
Citations: 162
h-index: 8
J. Balam
J. Balam
Citations: 985
h-index: 18

본 논문에서는 에이전트 기반 음성 인식을 위한 추론 전용 방식인 '보이스 메모리'를 제안합니다. 이 시스템에서, 고정된 수정 모듈은 스트림 처리 시간에 단일 도메인별 메모리 파일(.md)을 읽고, 각 발화에 대해 가설을 수정할지 아니면 생략하고 최상의 결과를 유지할지를 결정합니다. 비동기적으로 작동하는 점수 기반 최적화기는 제한적인 편집 작업을 통해 해당 파일을 수정하며, 수정이 별도의 평가 지표를 엄격하게 개선하는 경우에만 적용됩니다. 기존 ASR-LM 프레임워크에서 파생된 본 시스템은 '청취자-사고 모듈' 아키텍처라고 명명되며, 두 모듈은 메모리를 통해서만 연결되므로 가중치 변경이 없고, 학습된 기술은 감사 가능하며 이식성이 높습니다. 제약 조건은 이 루프가 발견하는 핵심적인 기술로 밝혀졌습니다. 제약 없이 작동하는 생성 오류 수정(GER) 방식은 금융 뉴스 데이터에서 최대 64%의 올바른 토큰을 잘못 수정하지만, 보이스 메모리는 이를 35%로 줄입니다. 'HyPoradise' 도메인 10개에 대해, 개방형 수정 모듈과 함께 사용할 때, 보이스 메모리는 가중 단어 오류율을 8.36%에서 7.52%로 낮추었으며 (세 개의 추가적인 컨텍스트 예제를 사용하면 7.47%), 모든 데이터셋의 기존 최상의 결과 성능을 저하시키지 않았습니다. 성능 향상은 항공 여행 관련 명령(8.40%에서 3.40%) 및 노이즈가 많은 원거리 음성(CHiME-4, 12.69%에서 10.46%)과 같은 영역에서 특히 두드러집니다. 이 메모리는 다양한 수정 모듈 간에 공유될 수 있으며, 추론 과정에 추가적인 파라미터를 생성하지 않습니다. 향후 연구를 위해 데모 및 예제 코드를 제공합니다.

Original Abstract

We present Voice Memory, a inference-only scheme for agentic speech recognition: at stream time, a frozen corrector reads a single per-domain memory.md and decides per utterance whether to act on the hypothesis or abstain and keep the 1-best. Asynchronously, a score-gated optimizer revises that file through bounded edits, accepting an edit only when it strictly improves a held-out score. Extended from classical ASR-LM framework, we refer this split the listener-thinker architecture; the two roles are coupled only through the memory, so no weights change and the learned skill stays auditable and portable. Restraint turns out to be the operative skill this loop discovers: unconstrained generative error correction (GER) over-corrects, breaking correct tokens on up to 64% of its edits on financial news, and Voice Memory, reduces this rate to 35%. Across ten HyPoradise domains with an open corrector, Voice Memory, lowers weighted word error rate from 8.36% to 7.52% (7.47% with three added in-context examples) without regressing any dataset below its 1-best baseline; gains concentrate where recoverable headroom is largest, including air-travel commands (8.40% to 3.40%) and noisy far-field speech (CHiME-4, 12.69% to 10.46%). The memory transfers across corrector families and adds zero parameters to the inference path. A demo and example code are provided for future studies.

0 Citations
0 Influential
9 Altmetric
45.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!