RW-Voice-EQ 벤치마크: 음성 AI 시스템 평가를 위한 실제 환경 기반 지표
RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems
현재의 음성 AI 벤치마크는 일반적으로 음성 명료도, 단어 오류율 또는 텍스트 기반 대화 품질과 같은 개별적인 기능만을 평가하며, 음성 데이터가 텍스트 표현과 구별되는 음향 정보를 시스템이 얼마나 활용하는지는 거의 평가하지 않습니다. 이에 따라 본 연구에서는 텍스트-음성 변환(TTS), 음성-음성 변환(STS), 음성 이해(SU) 및 자동 음성 인식(ASR)을 포괄하는 다차원적인 평가 지표인 '실제 환경 기반 음성 AI 벤치마크(Real World Voice EQ Bench)'를 제안합니다. 우리의 평가는 시스템 성능이 각 차원에 따라 크게 달라짐을 보여줍니다. TTS의 경우, 자연스러움, 표현력, 동일성 유지 및 신뢰성은 대체로 독립적인 평가 요소입니다. STS의 경우, 오디오 데이터에 접근하더라도 감정 표현이 사용되지 않는 경우가 있으며, 일부 에이전트는 여전히 주로 텍스트 기반으로 작동합니다. SU의 경우, 모델은 비언어적 작업에서 일관성 없는 성능을 보입니다. ASR의 경우, 실제 환경에서의 발음, 감정, 소음 및 대화 상황은 기존의 깨끗한 음성 데이터로 구성된 벤치마크에서는 파악하기 어려운 오류를 드러냅니다. 종합적으로 볼 때, 이러한 결과는 음성 AI가 음향적, 표현적, 상호 작용적 능력과 견고성이라는 다양한 측면을 갖춘 프로필로서 평가되어야 하며, 단일의 통합된 점수로 평가되어서는 안 된다는 것을 시사합니다.
Current voice AI benchmarks typically evaluate isolated capabilities such as speech intelligibility, word error rate, or text-based dialogue quality, but they rarely test whether systems harness the acoustic information that distinguishes spoken language from its textual representation. To this end, we introduce the Real World Voice EQ Bench, a multidimensional benchmark for evaluating voice AI across text-to-speech (TTS), speech-to-speech (STS), speech understanding (SU), and automatic speech recognition (ASR). Our evaluations indicate that performance is highly dimension-specific. For TTS, naturalness, expressiveness, identity stability, and reliability are largely independent evaluation dimensions. For STS, access to audio does not guarantee use of vocal affect, and some agents remain largely transcript-driven. For SU, models perform unevenly across paralinguistic tasks. For ASR, real world accent, emotion, noise, and conversational conditions expose failures that are not captured by established clean-speech benchmarks. Together, these results show that voice AI should be evaluated as a profile of acoustic, expressive, interactional, and robustness capabilities rather than by a single aggregate score.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.