SpeechEQ: 사회적 인식 능력을 갖춘 음성 대화 모델의 감정 지능 지수 벤치마킹
SpeechEQ: Benchmarking Emotional Intelligence Quotient in Socially Aware Voice Conversational Models
다중 모드 대화 시스템이 음성 상호작용을 점점 더 많이 활용함에 따라, 자연스러운 인간-AI 소통을 위한 핵심 과제가 파라링구이스틱(paralinguistic) 사회적 신호를 이해하는 능력입니다. 그러나 기존의 기계 감정 지능 평가 방법은 텍스트 또는 수동적인 음향 인식을 통해서만 추론 능력을 측정하며, 적극적인 다중 턴 대화에 필요한 복잡한 교차 모드 추론을 간과합니다. 본 논문에서는 Speech-Language Models (SLMs)의 사회언어적 추론 능력을 평가하기 위한 포괄적인 프레임워크인 extsc{SpeechEQ}를 소개합니다. 이 프레임워크는 EQ-i 2.0 이론에 기반한 15가지 감정 지수(EQ) 하위 요소로 구성된 검증된 2,265개의 대화 데이터 세트와 함께, 인간의 EQ 평가에서 영감을 받은 Spoken EQ (SEQ) 점수를 사용하여 다중 턴 평가 프로토콜을 제공합니다. 실험 결과, 기존의 음성 감정 인식 모델과 엔드-투-엔드 음성 언어 모델 모두가 음성을 통해 파라링구이스틱 신호를 이해하고 적용하는 데 한계가 있음을 보여줍니다. 엔드-투-엔드 아키텍처가 캐스케이드 시스템보다 성능이 좋지만, extsc{SpeechEQ}는 현재의 다중 모드 모델이 텍스트에 의존적인 "모달리티 단축(modality shortcut)", 정렬로 인한 "안전 함정(safety trap)" 및 "상황 기억 상실(contextual amnesia)"으로 인해 진정한 감정을 이해하는 AI로 발전하는 데 어려움을 겪고 있음을 보여줍니다. 본 논문의 벤치마크는 https://huggingface.co/datasets/SpeechEQ/SpeechEQ에서, 데모 페이지는 https://binomial14.github.io/speecheq-demo/ 에서 확인할 수 있습니다.
As multimodal conversational systems increasingly engage in spoken interaction, their ability to navigate paralinguistic social cues has become a critical bottleneck for natural human-AI communication. However, existing evaluations of machine emotional intelligence assess reasoning exclusively through isolated text or passive acoustic perception, overlooking the complex cross-modal reasoning required for active, multi-turn dialogue. We introduce \textsc{SpeechEQ}, a comprehensive framework designed to evaluate the sociolinguistic reasoning of Speech-Language Models (SLMs). The framework includes a validated dataset of 2,265 dialogues across 15 Emotional Quotient (EQ) subscales grounded in EQ-i 2.0 theory, along with a multi-turn evaluation protocol measured by our proposed Spoken EQ (SEQ) score inspired by human EQ assessments. Experiments show limitations in how both existing Speech Emotion Recognition and end-to-end Speech-Language Models understand and apply paralinguistic cues through speech. While end-to-end architectures outperform cascaded systems, \textsc{SpeechEQ} reveals that current multimodal models remain bottlenecked by a text-reliant ``modality shortcut,'' an alignment-induced ``safety trap,'' and ``contextual amnesia,'' highlighting the barriers to truly emotionally aware AI. Our benchmark can be accessed at https://huggingface.co/datasets/SpeechEQ/SpeechEQ and demo page at https://binomial14.github.io/speecheq-demo/
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.