2606.25529v1 Jun 24, 2026 cs.SD

STEB: 번역 정확성을 넘어 음성-음성 변환 시스템의 표현력 평가를 위한 벤치마크

STEB: A Speech-to-Speech Translation Expressiveness Benchmark for Evaluating Beyond Translation Fidelity

Songjun Cao
Songjun Cao
Citations: 288
h-index: 9
Long Ma
Long Ma
Citations: 27
h-index: 2
Sitong Cheng
Sitong Cheng
Citations: 195
h-index: 3
Weizhen Bian
Weizhen Bian
Citations: 209
h-index: 3
Bei Liu
Bei Liu
Citations: 82
h-index: 4
Chunyang Jiang
Chunyang Jiang
Citations: 25
h-index: 4
Yike Zhang
Yike Zhang
Citations: 177
h-index: 8
Weihao Wu
Weihao Wu
Citations: 7
h-index: 2
Yiming Li
Yiming Li
Citations: 33
h-index: 2
Chi-Min Chan
Chi-Min Chan
Citations: 3,363
h-index: 10
Wei Xue
Wei Xue
Citations: 133
h-index: 3
Jin Li
Jin Li
Citations: 55
h-index: 4

음성-음성 변환(S2ST)은 어휘적 의미뿐만 아니라 감정, 상황 스타일(예: 뉴스 보도와 극적인 대화), 그리고 비언어적 발성(NVs)과 같은 표현적 특성을 보존해야 합니다. 또한, 번역 정확도가 높으면서 원본 음성의 표현과 일관성이 있는 다국어 대상 음성을 대규모로 수집하는 것은 어렵기 때문에, 기준 기반 평가는 실용적이지 않습니다. 본 논문에서는 STEB(Speech-to-Speech Translation Expressiveness Benchmark)를 소개합니다. STEB는 32.6시간 분량의 중국어-영어 벤치마크 데이터셋으로, 표준적인 차원(번역 정확도, 화자 유사성, 시간 정렬)과 표현적 차원(감정, 상황 스타일, 비언어적 발성 보존)을 모두 평가합니다. 표현력 평가를 위해 STEB는 음성을 구조화된 표현적 속성으로 변환하고, LLM 심사관을 사용하여 원본 및 가설 속성을 비교하는 캡션-요약 프레임워크를 사용합니다. 인간 검증 결과, 모든 표현적 차원에서 청취자 판단과의 통계적으로 유의미한 상관관계가 나타났습니다. 본 논문에서는 다양한 S2ST 시스템(연쇄 시스템, 엔드 투 엔드 모델, 음성 대규모 언어 모델 포함)을 평가했습니다. 많은 시스템, 특히 연쇄 시스템은 뛰어난 번역 정확도를 달성했지만, 여전히 감정 보존(최고: 3.82/5) 및 비언어적 발성 보존(최고: 2.31/5)에 어려움을 겪습니다. 이러한 결과는 의미 전송과 표현 전송 간의 격차를 보여주며, 표현력 보존이 S2ST에서 해결해야 할 중요한 과제임을 시사합니다. 오디오 샘플은 https://cmots.github.io/steb.github.io/ 에서 확인할 수 있습니다.

Original Abstract

Speech-to-speech translation (S2ST) should preserve not only lexical meaning, but also expressive attributes: emotion, scenario style (e.g., news reporting vs. dramatic dialogue), and nonverbal vocalizations (NVs). Moreover, collecting cross-lingual target speech that is both translation-faithful and expressively aligned with the source is difficult at scale, making reference-based evaluation impractical. We introduce STEB (Speech-to-Speech Translation Expressiveness Benchmark), a 32.6-hour Chinese--English benchmark that evaluates both standard dimensions (translation fidelity, speaker similarity, duration alignment) and expressiveness dimensions (emotion, scenario style, NV preservation). For expressiveness evaluation, STEB uses a caption-then-summarize framework that converts speech into structured expressive attributes and compares source and hypothesis attributes with an LLM judge. Human validation shows statistically significant correlations with listener judgments across all expressive dimensions. We evaluate six S2ST systems covering cascaded systems, end-to-end models, and speech large language models. Many systems, especially cascaded ones, achieve strong translation fidelity, but they still struggle with emotion preservation (best: 3.82/5) and NV preservation (best: 2.31/5). These results reveal a gap between semantic transfer and expressive transfer, identifying expressiveness preservation as an open challenge for S2ST. Audio samples are available at https://cmots.github.io/steb.github.io/.

0 Citations
0 Influential
5 Altmetric
25.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!