2603.08936v1 Mar 09, 2026 cs.SD

VoxEmo: 음성 LLM을 활용한 음성 감정 인식 벤치마킹

VoxEmo: Benchmarking Speech Emotion Recognition with Speech LLMs

Hezhao Zhang
Hezhao Zhang
Citations: 84
h-index: 2
Huang-Cheng Chou
Huang-Cheng Chou
Citations: 9
h-index: 2
Shrikanth S. Narayanan
Shrikanth S. Narayanan
Citations: 10
h-index: 1
Thomas Hain
Thomas Hain
Citations: 24
h-index: 2

음성 대규모 언어 모델(LLM)은 생성형 인터페이스를 통해 음성 감정 인식(SER) 분야에서 큰 잠재력을 보여줍니다. 그러나 폐쇄형 분류에서 개방형 텍스트 생성을 사용하는 것은 제로샷(zero-shot) 방식의 불확실성을 야기하며, 이는 평가의 민감도를 프롬프트에 크게 의존하게 만듭니다. 또한, 기존의 음성 LLM 벤치마크는 인간 감정의 고유한 모호성을 간과하는 경향이 있습니다. 이에, 우리는 음성 LLM을 위한 포괄적인 SER 벤치마크인 VoxEmo를 제안합니다. VoxEmo는 15개 언어에 걸쳐 35개의 감정 데이터셋을 포함하며, 직접적인 분류부터 음성적 특징(paralinguistic) 추론에 이르기까지 다양한 프롬프트 복잡성을 갖춘 표준화된 도구를 제공합니다. 실제 인식 및 응용 환경을 반영하기 위해, 우리는 분포를 고려한 소프트 라벨(soft-label) 프로토콜과, 평가자 간의 의견 불일치를 모방하는 프롬프트 앙상블 전략을 도입했습니다. 실험 결과, 제로샷 음성 LLM은 하드 라벨 정확도 측면에서 지도 학습 기반 모델보다 낮은 성능을 보이지만, 인간의 주관적인 분포와 독특하게 일치하는 경향을 나타냅니다.

Original Abstract

Speech Large Language Models (LLMs) show great promise for speech emotion recognition (SER) via generative interfaces. However, shifting from closed-set classification to open text generation introduces zero-shot stochasticity, making evaluation highly sensitive to prompts. Additionally, conventional speech LLMs benchmarks overlook the inherent ambiguity of human emotion. Hence, we present VoxEmo, a comprehensive SER benchmark encompassing 35 emotion corpora across 15 languages for Speech LLMs. VoxEmo provides a standardized toolkit featuring varying prompt complexities, from direct classification to paralinguistic reasoning. To reflect real-world perception/application, we introduce a distribution-aware soft-label protocol and a prompt-ensemble strategy that emulates annotator disagreement. Experiments reveal that while zero-shot speech LLMs trail supervised baselines in hard-label accuracy, they uniquely align with human subjective distributions.

1 Citations
0 Influential
1 Altmetric
6.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!