2607.24430v1 Jul 27, 2026 cs.HC

당신을 자세히 살펴보겠습니다: 대화형 음성 합성 시스템을 위한 고급 표정 모델링

Let Me Look at You: Advanced Facial Expression Modeling for Conversational Speech Synthesis

Rui Liu
Rui Liu
Citations: 77
h-index: 6
Yifan Hu
Yifan Hu
IMU
Citations: 114
h-index: 6
Shuwei He
Shuwei He
Citations: 19
h-index: 2
Haizhou Li
Haizhou Li
Citations: 262
h-index: 8

대화형 음성 합성(Conversational Speech Synthesis, CSS)은 맥락에 적합하고 표현력이 뛰어나며 공감적인 음성을 생성하는 것을 목표로 하는 인간-컴퓨터 상호작용의 핵심 구성 요소입니다. 그러나 얼굴 표정은 공감적인 음성 상호 작용에 중요한 감정적 단서를 담고 있지만, 기존 접근 방식에서는 이러한 중요한 부분을 간과하는 경우가 많습니다. 또한, 음성과 시각 정보를 모두 포함하는 대규모 자연스러운 대화 데이터셋의 부족으로 인해 대화 환경에서의 시각적 감정 이해 기술 개발이 제한됩니다. 이러한 한계를 극복하기 위해, 본 연구에서는 대규모 언어 모델을 기반으로 구축된 표정 인식 CSS 프레임워크인 FacialTalker를 제안합니다. 얼굴 표정을 효율적으로 인코딩하기 위해, 각 프레임 단위의 얼굴 표정을 압축된 토큰으로 분산화하는 단일 코드북 시각 토크나이저인 AUTokenizer를 제안하며, 이는 얼굴 동작 단위(Action Units)의 조합을 통해 감독 학습됩니다. 또한, 시각 및 음성 토큰 시퀀스 모두에 대해 선호도 제약을 동시에 부과하여 모델의 표정 인식 능력과 다중 모달 대화 환경에서의 음성 의미 이해 능력을 향상시키는 이중 직접 선호도 최적화(DualDPO) 전략을 도입합니다. 더불어, 실제 인터넷 대화를 통해 완전히 자동화된 파이프라인으로 수집된 대규모 멀티모달 대화 데이터셋인 VSDD-1K를 구축했으며, 여기에는 85% 이상의 프레임에 유효한 얼굴이 포함된 1,033시간 이상의 동기화된 화자 비디오 및 음성 데이터가 포함되어 있습니다. 객관적 및 주관적 실험 결과는 FacialTalker가 표정 인식 능력과 음성 합성 품질 측면에서 강력한 기존 모델보다 일관되게 우수한 성능을 보이며, 더욱 자연스럽고 표현력이 풍부하며 대화 맥락에 더 잘 맞는 음성을 생성한다는 것을 보여줍니다. 이러한 결과는 또한 본 연구의 학습 전략 및 데이터셋 구축 파이프라인의 효과를 검증합니다.

Original Abstract

Conversational Speech Synthesis is a fundamental component of human-computer interaction, aiming to generate contextually appropriate, expressive, and empathetic speech. However, facial expressions encode subtle and rich affective cues that are crucial for empathetic speech interaction, whereas existing approaches often overlook this important modality. In addition, the lack of large-scale natural conversational datasets with both speech and visual modalities also limits the development of visual affect understanding in conversational settings.To address these limitations, we propose FacialTalker, a facial-expression-aware CSS framework built upon a large language model backbone. To efficiently encode facial expressions, we propose AUTokenizer, a single-codebook visual tokenizer that discretizes each frame-level facial expression into a compact token, trained with supervision from combinations of facial Action Units. We further introduce a dual direct preference optimization (DualDPO) strategy, which extends the DPO by jointly imposing preference constraints on both visual and speech token sequences, to enhance the model's understanding of facial expressions and speech semantics in multimodal conversational contexts. Moreover, we construct VSDD-1K, a large-scale multimodal dialogue dataset collected through a fully automated pipeline from real-world Internet conversations, comprising over 1,033 hours of synchronized speaker videos and speech, with more than 85\% of frames containing valid faces. Extensive objective and subjective experiments demonstrate that FacialTalker consistently outperforms strong baselines in facial-expression perception and speech synthesis quality, generating speech that is more natural, expressive, and better aligned with the conversational context. The results also validate the effectiveness of our training strategy and dataset construction pipeline.

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!