2603.15083v1 Mar 16, 2026 cs.CV

ReactMotion: 화자 발화를 기반으로 한 반응형 청취자 동작 생성

ReactMotion: Generating Reactive Listener Motions from Speaker Utterance

Ruibin Bai
Ruibin Bai
Citations: 62
h-index: 5
Rong Qu
Rong Qu
Citations: 59
h-index: 5
Bernard Ghanem
Bernard Ghanem
Citations: 42
h-index: 4
Cheng Luo
Cheng Luo
Citations: 755
h-index: 13
Bizhu Wu
Bizhu Wu
Citations: 294
h-index: 6
Bing Li
Bing Li
Citations: 71
h-index: 2
Jianfeng Ren
Jianfeng Ren
Citations: 73
h-index: 5
Linlin Shen
Linlin Shen
Citations: 255
h-index: 10

본 논문에서는 화자 발언에 적절하게 반응하는 자연스러운 청취자 신체 동작을 생성하는 새로운 과제인 '화자 발화를 기반으로 한 반응형 청취자 동작 생성'을 소개합니다. 그러나 이러한 비언어적 청취자 행동 모델링은 인간 반응의 본질적인 비결정적 특성으로 인해 아직 충분히 연구되지 않았으며 어려운 과제입니다. 이러한 과제를 해결하기 위해, 본 연구에서는 화자 발언과 다양한 적절성 정도를 가진 여러 후보 청취자 동작을 연결한 대규모 데이터셋인 ReactMotionNet을 제시합니다. 이 데이터셋은 청취자 행동의 일대다 관계를 명시적으로 반영하며, 단일의 정답 동작 외에 추가적인 정보를 제공합니다. 이러한 데이터셋 설계를 바탕으로, 입력-동작 정렬에 초점을 맞춘 기존의 동작 평가 지표를 고려하지 않고, 반응형 적절성을 평가하기 위한 선호도 기반 평가 프로토콜을 개발했습니다. 또한, 본 연구에서는 텍스트, 오디오, 감정, 그리고 동작을 통합적으로 모델링하고, 선호도 기반 학습 목표를 사용하여 적절하고 다양한 청취자 반응을 유도하는 통합 생성 프레임워크인 ReactMotion을 제안합니다. 광범위한 실험 결과, ReactMotion은 기존의 검색 기반 방법과 여러 LLM 기반 파이프라인보다 우수한 성능을 보이며, 더욱 자연스럽고, 다양하며, 적절한 청취자 동작을 생성하는 것으로 나타났습니다.

Original Abstract

In this paper, we introduce a new task, Reactive Listener Motion Generation from Speaker Utterance, which aims to generate naturalistic listener body motions that appropriately respond to a speaker's utterance. However, modeling such nonverbal listener behaviors remains underexplored and challenging due to the inherently non-deterministic nature of human reactions. To facilitate this task, we present ReactMotionNet, a large-scale dataset that pairs speaker utterances with multiple candidate listener motions annotated with varying degrees of appropriateness. This dataset design explicitly captures the one-to-many nature of listener behavior and provides supervision beyond a single ground-truth motion. Building on this dataset design, we develop preference-oriented evaluation protocols tailored to evaluate reactive appropriateness, where conventional motion metrics focusing on input-motion alignment ignore. We further propose ReactMotion, a unified generative framework that jointly models text, audio, emotion, and motion, and is trained with preference-based objectives to encourage both appropriate and diverse listener responses. Extensive experiments show that ReactMotion outperforms retrieval baselines and cascaded LLM-based pipelines, generating more natural, diverse, and appropriate listener motions.

0 Citations
0 Influential
6.5 Altmetric
32.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!