2608.03050v1 Aug 04, 2026 cs.SD

크로스 모달 부트스트래핑을 통한 피아노 편곡을 위한 음악 스타일 학습

Learning Music Style for Piano Arrangement Through Cross-Modal Bootstrapping

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Jingwei Zhao
Jingwei Zhao
Citations: 171
h-index: 6
Gus G. Xia
Gus G. Xia
Citations: 2,021
h-index: 25
Ye Wang
Ye Wang
Citations: 77
h-index: 5

음악 스타일이란 무엇인가? "스윙", "클래식", 또는 "감성적"과 같은 텍스트 레이블로 자주 설명되지만, 실제 스타일은 구체적인 음악 예시 안에 내재되어 있으며 숨겨져 있습니다. 본 논문에서는 원시 오디오 데이터로부터 암묵적인 음악 스타일을 학습하고 이를 심볼릭 음악 생성에 적용하는 크로스 모달 프레임워크를 소개합니다. BLIP-2에서 영감을 받아, 저희 모델은 쿼리 트랜스포머(Q-Former)를 사용하여 대규모 사전 훈련된 오디오 언어 모델(LM)로부터 스타일 표현을 추출하고, 이를 피아노 편곡 생성을 위한 심볼릭 LM에 적용합니다. 저희는 두 단계의 학습 전략을 채택합니다: 청각적 스타일과 심볼릭 표현을 정렬하기 위한 대비 학습, 그리고 음악 편곡을 위한 생성 모델링입니다. 저희 모델은 악보(내용)와 참조 오디오 예시(스타일) 모두를 조건으로 하여 피아노 연주를 생성하며, 이를 통해 제어 가능하고 스타일적으로 일관된 편곡이 가능합니다. 실험 결과는 저희 접근 방식이 피아노 커버 생성, 스타일 변환 및 오디오-MIDI 검색에 효과적임을 보여주며, 스타일 인지 정렬 및 음악 품질 측면에서 상당한 개선을 달성했습니다.

Original Abstract

What is music style? Though often described using text labels such as "swing," "classical," or "emotional," the real style remains implicit and hidden in concrete music examples. In this paper, we introduce a cross-modal framework that learns implicit music styles from raw audio and applies them to symbolic music generation. Inspired by BLIP-2, our model leverages a Querying Transformer (Q-Former) to extract style representations from a large, pre-trained audio language model (LM), and further applies them to condition a symbolic LM for generating piano arrangements. We adopt a two-stage training strategy: contrastive learning to align auditory style with symbolic expression, followed by generative modeling for music arrangement. Our model generates piano performances jointly conditioned on a lead sheet (content) and a reference audio example (style), enabling controllable and stylistically faithful arrangement. Experiments demonstrate the effectiveness of our approach in piano cover generation, style transfer, and audio-to-MIDI retrieval, achieving substantial improvements in style-aware alignment and music quality.

0 Citations
0 Influential
12.5 Altmetric
62.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!