2605.27984v1 May 27, 2026 cs.CL

KVoiceBench, KOpenAudioBench, 및 KMMAU: 음성 언어 모델 평가를 위한 에이전트 기반 한국어 음성 벤치마크

KVoiceBench, KOpenAudioBench, and KMMAU: Agent-Driven Korean Speech Benchmarks for Evaluating SpeechLMs

Haechan Kim
Haechan Kim
Citations: 8
h-index: 2
Seungjun Chung
Seungjun Chung
Citations: 109
h-index: 2
Inkyu Park
Inkyu Park
Citations: 58
h-index: 4
Jihoon Lee
Jihoon Lee
Citations: 0
h-index: 0
Jonghyun Lee
Jonghyun Lee
Citations: 6
h-index: 2

음성 언어 모델(SpeechLM)은 대규모 언어 모델(LLM)을 음성 모달리티로 확장하여 상당한 발전을 이루었습니다. 그러나 SpeechLM 평가는 여전히 영어에 편중되어 있어 다국어 음성 기능의 신뢰할 수 있는 평가를 제한합니다. 자동 음성 인식(ASR), 번역, 정규화 및 텍스트-음성 변환을 통한 단순한 벤치마크 이전은 언어별 지침, 답변 제약 조건 및 발화 형식을 손상시킬 수 있습니다. 또한 오디오 이해의 경우, 원본 언어 오디오를 이전하는 것은 대상 언어 화자의 특징, 억양 및 비언어적 특성을 보존하지 못합니다. 이러한 제한 사항을 해결하기 위해, 우리는 두 가지 인간-에이전트 기반 벤치마크 구축 프레임워크를 제안합니다. 하나는 원본 언어의 SpokenQA 벤치마크를 대상 언어의 SpokenQA 벤치마크로 변환하는 것이고, 다른 하나는 대상 언어의 ASR 코퍼스를 음성 인식(transcription) 및 화자 메타데이터를 사용하여 오디오 이해 벤치마크로 변환하는 것입니다. 이러한 프레임워크를 사용하여 한국어 SpokenQA를 위한 KVoiceBench와 KOpenAudioBench, 그리고 한국어 오디오 이해를 위한 KMMAU라는 세 가지 한국어 음성 벤치마크를 구축하고 공개적으로 배포했습니다. 총 12,345개의 샘플로 구성된 이 벤치마크를 사용하여 최근 개발된 여덟 개의 SpeechLM 모델을 평가한 결과, 영어-한국어 성능 격차는 모델 및 작업 유형에 따라 크게 달라지며, SpokenQA와 오디오 이해 순위가 다르게 나타나는 것을 확인했습니다. 이는 영어만으로 평가할 때 드러나지 않는 상호 보완적인 약점을 보여줍니다.

Original Abstract

Speech language models (SpeechLMs) have achieved substantial progress by extending large language models (LLMs) to the speech modality. However, SpeechLM evaluation remains heavily centered on English, limiting reliable assessment of multilingual speech capabilities. Straightforward benchmark transfer through ASR, translation, normalization, and TTS can corrupt language-specific instructions, answer constraints, and spoken forms; for audio understanding, transferring source-language audio also fails to preserve target-language speaker attributes, accents, and paralinguistic properties. To address these limitations, we propose two human-agent benchmark-construction frameworks: one transfers source-language SpokenQA benchmarks into target-language SpokenQA benchmarks, and the other converts target-language ASR corpora into audio understanding benchmarks using transcriptions and speaker metadata. Using these frameworks, we construct and publicly release three Korean speech benchmarks: KVoiceBench and KOpenAudioBench for Korean SpokenQA, and KMMAU for Korean audio understanding, comprising 12,345 samples in total. We evaluate eight recent SpeechLMs and find that English-Korean performance gaps vary substantially across models and task families, and that SpokenQA and audio understanding rankings diverge, revealing complementary weaknesses invisible to English-only evaluation.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!