2607.24013v2 Jul 27, 2026 cs.CV

AptAvatar: 고품질 오디오 기반 영상 생성 모델을 위한 빠르고 생생한 장편 아바타 제작 시스템

AptAvatar: Fast and Vivid Long-Form Audio-Driven Video Generation for Production-Ready Avatars

Meiguang Jin
Meiguang Jin
Citations: 3
h-index: 1
Junfeng Ma
Junfeng Ma
Citations: 3
h-index: 1
Hengyuan Zhang
Hengyuan Zhang
Citations: 228
h-index: 7
Jingna Sun
Jingna Sun
Citations: 0
h-index: 0

실용적인 수준의 오디오 기반 아바타 제작은 높은 품질과 표현력을 유지하면서 효율적인 추론이 가능해야 합니다. 그러나 기존 가속화 방법들은 종종 제한적인 구조 선택(예: 인과적 주의 메커니즘, 짧은 시간 범위) 또는 모델 용량 및 해상도 감소를 통해 품질을 희생합니다. 이러한 타협 없이, 저희는 140억 개의 파라미터를 가진 장편 오디오 기반 아바타 생성 프레임워크인 AptAvatar를 제안합니다. AptAvatar는 빠르고 풍부한 표현력을 제공하며, 실용적인 애플리케이션에서 효율성을 높이기 위해 극단적인 두 단계 생성 문제를 해결합니다. 다단계 모델과 두 단계 모델 간의 격차를 해소하기 위해, 저희는 Endpoint-Anchored Distribution Distillation을 도입했습니다. 이는 기존 분포 정합에 전용 Anchor Score Estimator를 추가하는 방식으로 작동하며, 이 Estimator는 사전 훈련된 4단계 중간 생성기에서 정의된 경로-종단 분포를 기반으로 학습됩니다. 이를 통해 진화하는 두 단계 모델은 달성 가능한 종단 레벨의 기준점을 얻을 수 있습니다. 또한, 장기간 일관성을 개선하기 위해 Self-Generated History Replay를 도입했습니다. 이는 조각 단위 훈련 중에 이전 생성기 체크포인트에서 캐시된 출력을 과거 조건으로 재사용하여 온라인 롤아웃 없이 추론 시점에 자체 생성된 과거 정보를 활용하는 효과를 내며, 누적된 과거 오류로 인한 품질 저하를 완화합니다. 광범위한 실험 결과, AptAvatar는 720p 해상도의 생생한 장편 아바타 영상을 단 2개의 Neural Feature Extraction(NFE) 연산으로 생성하며, 시각적 충실도를 유지하면서 60배의 속도 향상을 달성합니다. 관련 코드는 다음 링크에서 확인하실 수 있습니다: https://github.com/TaoLiveAIGC/AptAvatar

Original Abstract

Production-ready audio-driven avatar generation requires efficient inference without sacrificing fidelity or motion expressiveness. However, existing acceleration methods often compromise quality through restrictive architectural choices, such as causal attention and short temporal horizons, or by reducing model capacity and resolution. Without such compromises, we propose AptAvatar, a 14B-parameter long-form audio-driven avatar generation framework that delivers fast and expressive inference. For efficiency in production-level applications, AptAvatar addresses the extreme two-step generation challenge. To bridge the gap between the multi-step teacher model and the two-step student model, we introduce Endpoint-Anchored Distribution Distillation. It augments vanilla distribution matching with a dedicated Anchor Score Estimator trained on the trajectory-endpoint distribution defined from a frozen pretrained 4-step bridge generator. This provides an attainable endpoint-level anchor for the evolving two-step student. To improve long-horizon consistency, we further introduce Self-Generated History Replay, which reuses cached outputs from earlier generator checkpoints as history conditions during chunk-wise training. This approximates inference-time conditioning on self-generated histories without costly online rollouts, mitigating quality degradation from accumulated history errors. Extensive experiments demonstrate that AptAvatar generates vivid 720p long-form avatar videos with only 2 NFEs, achieving a 60x speedup while preserving visual fidelity and long-horizon identity. Code is available at https://github.com/TaoLiveAIGC/AptAvatar

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!