시각적 표현이 중요하다: 비디오-오디오 생성에서 시간적 차이를 활용
Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generation
비디오-오디오(V2A) 생성은 이미지-오디오(I2A) 생성을 확장하며, 오디오 합성에 필수적인 시간 정보를 제공하는 연속 프레임을 도입합니다. 그러나 기존의 조건부 확산 기반 V2A 방법은 일반적으로 추가적인 오디오-시각적 감독, 음향 구조 예측 또는 대규모 멀티모달 모델을 활용하여 시각적 조건을 강화하는데, 이는 추가 네트워크나 강력한 유도 편향을 필요로 합니다. 최근의 시각적 표현 학습 연구를 바탕으로, 우리는 시간적 차이(TD)를 V2A와 I2A를 구분하는 핵심 표현으로 활용하는 TD-V2A를 제안합니다. TD는 최소한의 아키텍처 변경만으로 시각적 조건을 풍부하게 합니다. 먼저, 프레임 및 특징 수준 모두에서 TD를 조사하여 TD가 시각적 표현을 어떻게 보완하는 데 가장 효과적인지 파악합니다. 이러한 결과를 바탕으로, 계층적 지속 학습 전략과 점진적으로 감소하는 시간적 차이 안내 방법을 개발하여 확산 훈련 및 샘플링 과정에서 TD 정보를 점진적으로 학습하고 활용합니다. 표준 데이터 세트에서의 광범위한 실험 결과는 제안된 프레임워크를 통해 TD를 효과적으로 활용하면 엔드-투-엔드 V2A 생성 품질을 크게 향상시킬 수 있으며, 심지어 대비적 오디오-시각 사전 훈련과 같은 특수하게 설계된 V2A 표현보다도 우수한 성능을 보인다는 것을 보여줍니다.
Video-to-audio (V2A) generation extends image-to-audio generation (I2A) by introducing consecutive frames that provide essential temporal cues for audio synthesis. However, existing conditional diffusion-based V2A methods typically enhance visual conditioning with additional audio-visual supervision, acoustic structure prediction, or reasoning from large multimodal models, requiring extra networks or strong inductive biases. Inspired by recent advances in visual representation learning, we introduce TD-V2A, which leverages temporal differences (TD) as the key representation that distinguishes V2A from I2A, enriching visual conditioning with minimal architectural modification. We first investigate TD at both the frame and feature levels to identify the most effective representation level at which TD complements visual representations. Based on these findings, we develop a hierarchically continual learning strategy and an annealed temporal differences guidance method to progressively learn and exploit TD information during diffusion training and sampling process, respectively. Extensive experiments on benchmark datasets demonstrate that effectively exploiting TD through our proposed framework significantly improves end-to-end V2A generation quality, even outperforming dedicated V2A representations such as contrastive audio-visual pretraining.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.