2607.05365v1 Jul 06, 2026 cs.CL

SPEARBench: 스트리밍 음성-음성 언어 모델의 자연스러움 평가를 위한 벤치마크

SPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language Models

Thomas Thebaud
Thomas Thebaud
Citations: 112
h-index: 6
L. Moro-Velázquez
L. Moro-Velázquez
Citations: 1,718
h-index: 24
Yuzhe Wang
Yuzhe Wang
Citations: 13
h-index: 2
Sathvik Manikantan Napa Ugandhar
Sathvik Manikantan Napa Ugandhar
Citations: 0
h-index: 0
Hao Zhang
Hao Zhang
Citations: 0
h-index: 0
Georgi Tinchev
Georgi Tinchev
Citations: 289
h-index: 8
Venkatesh Ravichandran
Venkatesh Ravichandran
Citations: 8
h-index: 1
Ashish G. Hallur
Ashish G. Hallur
Citations: 14
h-index: 1

스트리밍 음성-음성 언어 모델은 구두 질문에 대해 합성 음성을 통해 직접 응답하는 것을 목표로 합니다. 그러나 기존의 음성 및 텍스트 벤치마크는 이러한 시스템이 대화에서 얼마나 자연스럽게 작동하는지를 제대로 반영하지 못합니다. 대화에서의 자연스러움은 타이밍, 발언 교대, 운율, 상호작용 방식, 언어 및 방언 일관성, 관계에 대한 적절성과 같은 다양한 요소들이 복합적으로 작용하여 결정됩니다. 본 논문에서는 질문-응답 상호작용을 통해 음성-음성 언어 모델의 자연스러움을 평가하기 위한 벤치마크인 SPEARBench를 소개합니다. SPEARBench는 Seamless Interaction 코퍼스를 기반으로 제어된 대화 프롬프트를 구성하고, 여러 모델에 대해 추론을 수행하며, 응답 지연 시간, 중단 현상, 음성 품질, ASR (자동 음성 인식) 강건성, 언어 및 방언 일관성, 감정적 자연스러움, 상호작용 방식, 그리고 설명 가능한 기준선과 같은 다차원적인 프로토콜을 사용하여 생성된 답변을 평가합니다. 벤치마크에는 원본 인간의 답변이 참조 조건으로 포함되어 있으며, 여러 최신 모델에 대한 결과를 보고합니다. 결과는 현재 모델들이 높은 수준의 음성 품질과 낮은 ASR 오류를 달성할 수 있지만, 지연 시간, 중복 현상, 방언 유지, 감정적 적응 및 상호작용 방식 측면에서 인간의 대화 행동과 여전히 차이가 있음을 보여줍니다.

Original Abstract

Streaming speech-to-speech language models aim to answer spoken queries directly with synthetic speech. However, standard speech and text benchmarks do not capture whether these systems behave naturally in conversations, where timing, turn-taking, prosody, interpersonal stance, language and dialect consistency, and relationship-aware appropriateness jointly shape perceived quality. We introduce SPEARBench, a benchmark for evaluating naturalness in speech-to-speech language models from question-answer interactions. SPEARBench constructs controlled dialogue prompts from the Seamless Interaction corpus, runs inference across multiple models, and evaluates generated answers using a multidimensional protocol that covers response latency, interruptions, speech quality, ASR robustness, language and dialect consistency, emotional naturalness, interpersonal stance, and explainable distributional baselines. The benchmark includes original human answers as a reference condition and reports results for several contemporary models. Results show that current models can achieve high signal-level quality and low ASR error while still differing from human conversational behavior in latency, overlap, dialect preservation, emotional adaptation, and interpersonal stance dynamics.

0 Citations
0 Influential
12 Altmetric
60.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!