2604.14493v2 Apr 16, 2026 cs.AI

온장치 스트리밍 음성 인식의 한계를 뛰어넘기: 저지연 추론을 위한 작고 고정밀 영어 모델

Pushing the Limits of On-Device Streaming ASR: A Compact, High-Accuracy English Model for Low-Latency Inference

Nenad Banfic
Nenad Banfic
Citations: 1
h-index: 1
D. Fan
D. Fan
Citations: 110
h-index: 6
K. Vaishnavi
K. Vaishnavi
Citations: 1
h-index: 1
S. Kemp
S. Kemp
Citations: 1
h-index: 1
Ruifeng Ren
Ruifeng Ren
Citations: 58
h-index: 5
S. Shaw
S. Shaw
Citations: 28
h-index: 2
Meng Tang
Meng Tang
Citations: 38
h-index: 3
S. Choi
S. Choi
Citations: 5
h-index: 2

고품질 자동 음성 인식(ASR)을 엣지 장치에 구현하려면 정확도, 지연 시간, 메모리 사용량을 동시에 최적화하고 GPU 가속 없이 CPU에서만 작동하는 모델이 필요합니다. 본 연구에서는 인코더-디코더, 트랜스듀서, LLM 기반 아키텍처를 포함한 최첨단 ASR 아키텍처에 대한 체계적인 실험적 연구를 수행하고, 배치, 청킹, 스트리밍 추론 모드를 기준으로 평가했습니다. OpenAI Whisper, NVIDIA Nemotron, Parakeet TDT, Canary, Conformer Transducer, Qwen3-ASR을 포함한 50개 이상의 구성에 대한 종합적인 벤치마크를 통해, NVIDIA의 Nemotron Speech Streaming이 제한된 리소스 하드웨어에서 실시간 영어 스트리밍에 가장 적합한 후보임을 확인했습니다. 이후 전체 스트리밍 추론 파이프라인을 ONNX Runtime으로 재구현하고, 중요도 기반 k-양자화, 혼합 정밀도 방식, 근사값 반올림 양자화 등 다양한 사후 양자화 전략을 적용하여 그래프 수준의 연산자 융합과 함께 제어된 평가를 수행했습니다. 이러한 최적화 과정을 통해 모델 크기를 2.47GB에서 0.67GB까지 줄이면서, 단어 오류율(WER)을 전체 정밀도 PyTorch 기준의 1% 이내로 유지했습니다. 권장 구성인 int4 k-양자화 변형은 8가지 표준 벤치마크에서 평균 스트리밍 WER 8.20%를 달성하며, 0.56초의 알고리즘 지연 시간으로 CPU에서 실시간보다 빠르게 실행되어, 온장치 스트리밍 ASR의 새로운 품질-효율 균형점을 제시합니다.

Original Abstract

Deploying high-quality automatic speech recognition (ASR) on edge devices requires models that jointly optimize accuracy, latency, and memory footprint while operating entirely on CPU without GPU acceleration. We conduct a systematic empirical study of state-of-the-art ASR architectures, encompassing encoder-decoder, transducer, and LLM-based paradigms, evaluated across batch, chunked, and streaming inference modes. Through a comprehensive benchmark of over 50 configurations spanning OpenAI Whisper, NVIDIA Nemotron, Parakeet TDT, Canary, Conformer Transducer, and Qwen3-ASR, we identify NVIDIA's Nemotron Speech Streaming as the strongest candidate for real-time English streaming on resource-constrained hardware. We then re-implement the complete streaming inference pipeline in ONNX Runtime and conduct a controlled evaluation of multiple post-training quantization strategies, including importance-weighted k-quant, mixed-precision schemes, and round-to-nearest quantization, combined with graph-level operator fusion. These optimizations reduce the model from 2.47 GB to as little as 0.67 GB while maintaining word error rate (WER) within 1% absolute of the full-precision PyTorch baseline. Our recommended configuration, the int4 k-quant variant, achieves 8.20% average streaming WER across eight standard benchmarks, running comfortably faster than real-time on CPU with 0.56 s algorithmic latency, establishing a new quality-efficiency Pareto point for on-device streaming ASR.

2 Citations
1 Influential
3 Altmetric
19.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!