아바타 포싱: 자연스러운 대화를 위한 실시간 인터랙티브 헤드 아바타 생성
Avatar Forcing: Real-Time Interactive Head Avatar Generation for Natural Conversation
대화형 헤드 생성 기술은 정적인 초상화를 기반으로 사실적인 아바타를 만들어 가상 커뮤니케이션 및 콘텐츠 제작에 활용됩니다. 그러나 현재 모델은 진정한 인터랙티브 커뮤니케이션의 느낌을 제대로 전달하지 못하며, 종종 감정적인 참여가 부족한 일방적인 응답을 생성하는 경우가 많습니다. 우리는 진정한 인터랙티브 아바타를 구현하기 위한 두 가지 주요 과제를 제시합니다. 첫째, 인과 관계 제약 조건 하에서 실시간으로 움직임을 생성하는 것, 둘째, 추가적인 레이블 데이터 없이도 표현력이 풍부하고 생동감 넘치는 반응을 학습하는 것입니다. 이러한 과제들을 해결하기 위해, 우리는 실시간 사용자-아바타 상호 작용을 확산 모델링을 통해 표현하는 새로운 프레임워크인 '아바타 포싱'을 제안합니다. 이 설계는 아바타가 사용자의 음성 및 움직임과 같은 실시간 멀티모달 입력을 낮은 지연 시간으로 처리하여, 음성, 끄덕임, 웃음과 같은 언어적 및 비언어적 신호에 즉각적으로 반응할 수 있도록 합니다. 또한, 사용자 조건을 제거하여 생성된 합성적인 부정 샘플을 활용하는 직접적인 선호도 최적화 방법을 도입하여, 레이블이 없는 환경에서도 표현력 있는 상호 작용을 학습할 수 있도록 합니다. 실험 결과는 우리의 프레임워크가 약 500ms의 낮은 지연 시간으로 실시간 상호 작용을 가능하게 하며, 기준 모델 대비 6.8배의 속도 향상을 달성하고, 반응적이고 표현력이 풍부한 아바타 움직임을 생성하여, 기준 모델보다 80% 이상의 선호도를 얻는다는 것을 보여줍니다.
Talking head generation creates lifelike avatars from static portraits for virtual communication and content creation. However, current models do not yet convey the feeling of truly interactive communication, often generating one-way responses that lack emotional engagement. We identify two key challenges toward truly interactive avatars: generating motion in real-time under causal constraints and learning expressive, vibrant reactions without additional labeled data. To address these challenges, we propose Avatar Forcing, a new framework for interactive head avatar generation that models real-time user-avatar interactions through diffusion forcing. This design allows the avatar to process real-time multimodal inputs, including the user's audio and motion, with low latency for instant reactions to both verbal and non-verbal cues such as speech, nods, and laughter. Furthermore, we introduce a direct preference optimization method that leverages synthetic losing samples constructed by dropping user conditions, enabling label-free learning of expressive interaction. Experimental results demonstrate that our framework enables real-time interaction with low latency (approximately 500ms), achieving 6.8X speedup compared to the baseline, and produces reactive and expressive avatar motion, which is preferred over 80% against the baseline.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.