2608.02474v1 Aug 03, 2026 cs.CV

EchoCache: 에너지 기반의 교차 모드 캐싱을 통한 효율적인 오디오 기반 비디오 생성

EchoCache: Energy-Guided Cross-Modal Caching for Efficient Audio-Driven Video Generation

Jiayu Chen
Jiayu Chen
Citations: 41
h-index: 4
Zihao Zheng
Zihao Zheng
Citations: 34
h-index: 3
Maoliang Li
Maoliang Li
Citations: 46
h-index: 4
Xiang Chen
Xiang Chen
Citations: 49
h-index: 4
Guojie Luo
Guojie Luo
Citations: 19
h-index: 3
H. Zou
H. Zou
Citations: 11
h-index: 3
Xiaoyu Wu
Xiaoyu Wu
Citations: 0
h-index: 0
Rongshan Gao
Rongshan Gao
Citations: 0
h-index: 0
Xinhao Sun
Xinhao Sun
Citations: 38
h-index: 3

오디오 기반 비디오 생성(A2V)은 시간적으로 일관되고 오디오-시각적으로 정렬된 비디오를 합성하는 데 있어 상당한 발전을 이루었지만, 확산 모델의 반복적인 노이즈 제거 과정으로 인해 추론 비용이 여전히 높습니다. 기존 캐싱 방법은 주로 시각적 특징에서의 시간적 중복성을 활용하지만, A2V에서 오디오가 시각 생성에 영향을 미치는 방식은 시간적으로 불균일한 중요도를 가지므로 교차 모드 정렬을 간과합니다. 본 논문에서는 기존 A2V 캐싱 방법에서 발견되는 두 가지 유형의 불일치(시간-의미론적 불일치 및 계산-저장소 불일치)를 제시하고, 이를 해결하기 위해 에너지 기반의 교차 모드 캐싱 프레임워크인 EchoCache를 제안합니다. EchoCache는 오디오의 시간-주파수 에너지를 중요한 지표로 활용하여 잠재 수준에서의 캐시 업데이트를 안내하며, 효율성과 메모리 최적화를 동시에 달성하기 위해 양자화된 캐시 관리를 포함하는 동적인 타임스텝-잠재 캐싱 메커니즘을 도입합니다. 주류 A2V 모델에 대한 광범위한 실험 결과, EchoCache는 생성 품질과 오디오-시각적 일관성을 유지하면서도 지연 시간과 품질 간의 균형을 지속적으로 향상시키는 것으로 나타났습니다. 특히, EMTD 벤치마크에서 Wan2.2-S2V 모델에 대해 EchoCache는 2.46배의 속도 향상을 달성했으며, 전반적인 성능이 가장 우수했습니다. 코드 및 관련 자료는 https://github.com/IF-LAB-PKU/EchoCache 에서 확인할 수 있습니다.

Original Abstract

Audio-driven video generation (A2V) has achieved promising progress in synthesizing temporally coherent and audio-visually aligned videos, yet its inference remains expensive due to the iterative denoising process of diffusion models. Existing caching methods mainly exploit temporal redundancy in visual features while overlooking the cross-modal alignment of A2V, where audio drives visual generation with highly non-uniform temporal importance. In this paper, we identify two levels of misalignment in existing A2V caching methods: temporal-semantic and computation-storage misalignment. To address them, we propose EchoCache, an energy-guided cross-modal caching framework for efficient A2V generation. EchoCache leverages audio time-frequency energy as a saliency anchor to guide latent-level cache updates and further introduces a dynamic timestep-latent caching mechanism with quantized cache management for joint efficiency and memory optimization. Extensive experiments on mainstream A2V models show that EchoCache consistently improves the latency-quality trade-off while preserving generation quality and audio-visual consistency. In particular, on Wan2.2-S2V over the EMTD benchmark, EchoCache achieves a 2.46x speedup with the best overall performance. Code is available at https://github.com/IF-LAB-PKU/EchoCache.

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!