2608.00079v1 Jul 29, 2026 cs.CV

LeapTalk: 음성 합성 기술에서 지연 시간과 품질 간의 균형을 깨다

LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generation

Songhua Liu
Songhua Liu
Citations: 761
h-index: 12
Rongxiang Zhang
Rongxiang Zhang
Citations: 0
h-index: 0

장시간 및 실시간 음성 합성 기술은 지연 시간과 품질 사이의 상충 관계 때문에 여전히 어려운 과제입니다. 비효율적인 다단계 확산 모델은 스트리밍 생성을 방해하는 반면, 실시간 순환적 접근 방식은 오류 누적과 얼굴 특징 변화 문제를 야기합니다. 이러한 단점을 해결하기 위해, 우리는 LeapTalk이라는 새로운 프레임워크를 제안합니다. LeapTalk은 단일 단계로 안정적이고 실시간 음성 합성 기술을 구현하며, 임의의 길이의 영상으로 확장 가능합니다. 우리 접근 방식의 핵심은 단일 단계 브리지 증류(bridge distillation) 기법입니다. 기존의 노이즈-데이터 패러다임에서 벗어나, 우리는 브라운 운동 기반의 데이터-데이터 전송 방식을 도입했습니다. 지속적인 참조점을 사용하여 이 전략은 얼굴 특징 변화를 효과적으로 완화하고 장기적인 시간적 안정성을 향상시킵니다. 또한, 사전 훈련된 확산 모델(선생님 모델)에서 학생 브리지 모델로 원활한 지식 전달을 가능하게 하기 위해, SNR 정렬 기반의 시간 변환 $Φ(τ)$을 사용하는 이질적 증류 프레임워크를 탐구했습니다. 더욱이, 극단적인 단계 감소 상황에서도 미세한 입술 동기화를 유지하기 위해 오디오 기반의 분류기-프리 가이드(classifier-free guidance) 메커니즘을 제안합니다. 광범위한 실험 결과는 우리 방법이 단 1단계로 최대 200 FPS까지 고품질의 시간적으로 일관된 영상을 생성하며, 효율성과 안정성 측면에서 기존 방식보다 훨씬 우수한 성능을 보임을 보여줍니다. 프로젝트 페이지: https://zhangrongxiang.github.io/leaptalk-page/

Original Abstract

Long-form and real-time talking-head generation remains challenging due to a latency-quality trade-off: inefficient multi-step diffusion prohibits streaming generation, whereas real-time autoregressive approaches suffer from error accumulation and identity drift. To address this drawback, we propose LeapTalk, a novel framework that achieves stable and real-time talking-head generation with a single forward step, scaling to arbitrarily long videos. At the heart of our approach lies a single-step bridge distillation scheme. On the one hand, departing from the conventional noise-to-data paradigm, we introduce a data-to-data transport formulation based on a Brownian bridge. Anchored by a persistent reference, this strategy effectively mitigates identity drift and enhances long-term temporal stability. On the other hand, to enable smooth knowledge transfer from a pre-trained diffusion teacher to the student bridge model, we explore a heterogeneous distillation framework with an SNR-aligned time transformation $Φ(τ)$, which bridges the functional discrepancy between the two models. Moreover, we propose an audio-driven classifier-free guidance mechanism to maintain fine-grained lip synchronization under extreme step reduction. Extensive experiments demonstrate that our method achieves high-fidelity and temporally consistent video generation with only 1 step at up to 200 FPS, significantly outperforming existing approaches in both efficiency and stability. Project Page: https://zhangrongxiang.github.io/leaptalk-page/

0 Citations
0 Influential
6 Altmetric
30.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!