2607.24359v1 Jul 27, 2026 cs.CV

TaoMate: 앵커 기반 메모리 브리지 - 실시간 오디오-비디오 디지털 휴먼 생성 시 진화하는 상태와 참조 상태 간의 연결

TaoMate: Anchor-Guided Memory Bridging Evolving and Reference States for Real-Time Audio-Video Digital Human Generation

Qijun Gan
Qijun Gan
Citations: 136
h-index: 5
Chenwei Zhang
Chenwei Zhang
Citations: 0
h-index: 0
Meiguang Jin
Meiguang Jin
Citations: 3
h-index: 1
Junfeng Ma
Junfeng Ma
Citations: 3
h-index: 1
Qiu Shen
Qiu Shen
Citations: 0
h-index: 0

실시간으로 긴 형식의 디지털 휴먼 생성을 위해서는 원활한 오디오-비주얼 콘텐츠 확장을 위해 인과 모델이 필요하며, 이를 통해 주체의 외형을 유지하고 연속적인 세그먼트 전반에 걸쳐 오디오-비디오 동기화를 보장해야 합니다. 제한된 캐시는 로컬 동작 및 음성 맥락을 유지하지만 오래된 정보를 버리는 반면, 전체 생성 이력을 참조하는 것은 계산 비용이 많이 들고 누적된 오류를 야기할 수 있습니다. 본 논문에서는 exttt{TaoMate}, 즉 앵커 기반의 지속적인 메모리 프레임워크를 제안합니다. 이 프레임워크는 불변(immutable)인 시각적 앵커를 유지하고, 완성된 비디오 및 오디오 블록을 고정 용량의 동적 상태로 압축하며, 활성 캐시를 확장하지 않고 모달리티별 잔차 어텐션을 통해 이러한 상태를 검색합니다. 또한 참조 정보를 인식하는 변조 방법을 사용하여 동적 및 앵커 외형 통계에 기반하여 비디오 특징을 조정합니다. 앵커를 보존하면서 인과적 맥락을 추출하는 기술은 롤아웃 호라이즌, 접두사 출처, 캐시-히스토리 신뢰성을 다양하게 변화시키면서도 불변의 시각적 앵커가 변경되지 않도록 합니다. exttt{TaoMate}는 지속적인 메모리를 스테이지별 노이즈 제거 의존성으로부터 분리하여 스테이지 병렬 처리를 가능하게 하며, 파이프라인 특유의 재학습 없이 자기 회귀 추론을 가속화합니다. 우리는 외형, 시간적 일관성, 동기화, 얼굴 특징 및 음성 품질에 대한 진단 평가를 통해 긴 형식의 비디오 연속 생성을 평가했습니다. 결과는 exttt{TaoMate}가 프롬프트 기반 세그먼트에서 안정적인 외형을 유지하고 자기 회귀 생성 과정에서 강력한 오디오-비주얼 동기화를 달성한다는 것을 보여줍니다. 프로젝트 페이지는 https://taoliveaigc.github.io/TaoMate 입니다.

Original Abstract

Real-time long-form digital-human generation relies on causal models to extend audio-visual content while preserving subject appearance and audio-video synchronization across successive segments. A bounded cache retains local motion and phonetic context but discards older evidence, whereas attending to the complete generated history is computationally expensive and can propagate accumulated errors. We present \method, an anchor-guided persistent-memory framework for few-step joint audio-video generation. The framework preserves an immutable visual anchor, compresses completed video and audio blocks into fixed-capacity dynamic states, and retrieves those states through modality-specific residual attention without extending the active cache. A reference-aware modulation method additionally conditions video features on dynamic and anchor appearance statistics. Anchor-preserving causal-context distillation varies rollout horizon, prefix provenance, and cache-history reliability while keeping the immutable visual anchor unperturbed. By separating persistent memory from stage-local denoising dependencies, \method further admits stage-parallel execution across blocks, accelerating autoregressive inference without pipeline-specific retraining. We evaluate long-form video continuations with appearance, temporal, synchronization, facial, and speech diagnostics. Results show that \method preserves stable appearance across prompt-conditioned segments and strong audio-visual synchronization under autoregressive generation. Our project page is https://taoliveaigc.github.io/TaoMate.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!