V2M-Zero: 제로-페어 시간 정렬 비디오-음악 생성
V2M-Zero: Zero-Pair Time-Aligned Video-to-Music Generation
기존의 텍스트-음악 모델은 세밀한 시간 제어가 부족하여 비디오 이벤트와 시간적으로 일치하는 음악을 생성하는 데 어려움이 있습니다. 본 연구에서는 비디오에 맞춰 시간적으로 정렬된 음악을 생성하는 제로-페어 방식인 V2M-Zero를 소개합니다. 저희 방법은 중요한 관찰에서 비롯되었습니다. 즉, 시간 동기화는 언제, 얼마나 변화가 발생하는지를 매칭하는 것이 중요하며, 어떤 변화가 발생하는지는 중요하지 않습니다. 음악적 및 시각적 이벤트는 의미적으로 다르지만, 각 모달리티 내에서 독립적으로 캡처될 수 있는 공유된 시간 구조를 가지고 있습니다. 저희는 사전 훈련된 음악 및 비디오 인코더를 사용하여 계산된 이벤트 곡을 통해 이러한 구조를 캡처합니다. 각 모달리티 내에서 시간적 변화를 독립적으로 측정함으로써, 이러한 곡은 모달리티 간 비교 가능한 표현을 제공합니다. 이를 통해 간단한 훈련 전략을 사용할 수 있습니다. 즉, 텍스트-음악 모델을 음악 이벤트 곡으로 미세 조정하고, 훈련 과정에서 크로스-모달 훈련이나 페어링된 데이터 없이 추론 시 비디오 이벤트 곡을 대체합니다. OES-Pub, MovieGenBench-Music, AIST++ 데이터셋에서 V2M-Zero는 페어링된 데이터 기반 모델보다 5-21% 높은 음질, 13-15% 더 나은 의미적 정렬, 21-52% 향상된 시간 동기화, 그리고 댄스 비디오에서 28% 더 높은 비트 정렬을 달성했습니다. 대규모 크라우드 소싱 주관 청취 테스트에서도 유사한 결과를 얻었습니다. 종합적으로, 저희의 결과는 시간 정렬이 페어링된 크로스-모달 감독이 아닌, 모달리티 내의 특징을 통해 이루어질 때 비디오-음악 생성에 효과적이라는 것을 입증합니다. 결과는 https://genjib.github.io/v2m_zero/ 에서 확인할 수 있습니다.
Generating music that temporally aligns with video events is challenging for existing text-to-music models, which lack fine-grained temporal control. We introduce V2M-Zero, a zero-pair video-to-music generation approach that outputs time-aligned music for video. Our method is motivated by a key observation: temporal synchronization requires matching when and how much change occurs, not what changes. While musical and visual events differ semantically, they exhibit shared temporal structure that can be captured independently within each modality. We capture this structure through event curves computed from intra-modal similarity using pretrained music and video encoders. By measuring temporal change within each modality independently, these curves provide comparable representations across modalities. This enables a simple training strategy: fine-tune a text-to-music model on music-event curves, then substitute video-event curves at inference without cross-modal training or paired data. Across OES-Pub, MovieGenBench-Music, and AIST++, V2M-Zero achieves substantial gains over paired-data baselines: 5-21% higher audio quality, 13-15% better semantic alignment, 21-52% improved temporal synchronization, and 28% higher beat alignment on dance videos. We find similar results via a large crowd-source subjective listening test. Overall, our results validate that temporal alignment through within-modality features, rather than paired cross-modal supervision, is effective for video-to-music generation. Results are available at https://genjib.github.io/v2m_zero/
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.