전체 곡 생성의 한계를 뛰어넘는 연구: 계층적 자기 회귀 계획과 플로우 매칭 렌더링의 결합
Pushing the Frontier of Full-Song Generation: Hierarchical Autoregressive Planning Meets Flow-Matching Rendering
본 논문에서는 가사, 텍스트 설명 및 음악적 속성을 기반으로 고품질의 전체 길이 음악을 생성할 수 있는 통합된 곡 생성 프레임워크를 제시합니다. 제안하는 프레임워크는 세 가지 작업을 지원합니다: 텍스트 설명을 기반으로 한 완전한 노래 생성 (Lyrics-to-Song Generation), 보컬이 없는 음악 생성 (Instrumental Music Generation), 그리고 기존 노래의 스타일을 새롭게 해석하면서 멜로디 내용을 유지하는 커버곡 생성 (Cover Song Generation). 시스템 아키텍처는 의미 정보를 고려한 토크나이저, 하이브리드 언어 모델 (hybird-LM), FullDiT 및 두 단계 멜로디 모듈의 네 가지 주요 구성 요소로 구성됩니다. 토크나이저는 오디오를 효율적인 이산 음악 표현을 위한 8-코드북 RVQ 토큰으로 인코딩합니다. 이러한 토큰을 기반으로, hybird-LM은 전체 곡 생성을 위해 계층적 자기 회귀 방식으로 오디오 토큰 모델링을 수행합니다. 오디오 품질 향상을 위해, FullDiT는 코덱 토큰, 가사 및 텍스트 설명을 조건으로 하는 연속적인 VAE 잠재 공간에서 전체 곡의 플로우 매칭을 수행합니다. 커버곡 생성의 경우, 멜로디 모듈은 참조 오디오에서 멜로디 정보를 추출하고 이산화하여 생성 과정을 안내하면서 원래 멜로디 내용을 유지합니다. 마지막으로, hybird-LM에 대한 보상 기반 후 학습 전략으로 DPO, GRPO 및 OPD를 조사하고, FullDiT에 플로우 기반의 GRPO를 적용하여 음악성과 렌더링 품질을 향상시켰습니다. 다국어 자동 평가 벤치마크와 Artificial Analysis Music with Vocals 리더보드를 활용한 실험 결과는 제안하는 프레임워크가 다양한 환경에서 경쟁력 있는 성능을 달성함을 보여줍니다.
In this report, we present a unified song generation framework capable of producing high-quality full-length music from lyrics, text descriptions, and musical attributes. The proposed framework supports three tasks: Lyrics-to-Song Generation, which generates complete songs from text descriptions, lyrics, and musical attributes; Instrumental Music Generation, which creates music without vocals; and Cover Song Generation, which reinterprets existing songs with different styles while preserving their melodic content. Architecturally, our system consists of four main components: a semantic-aware tokenizer, hybird-LM, FullDiT, and a two-level melody module. The tokenizer encodes audio into 8-codebook RVQ tokens for efficient discrete music representation. Based on these tokens, hybird-LM performs hierarchical autoregressive audio-token modeling for full-song generation. To improve audio fidelity, FullDiT performs full-song flow matching in a continuous VAE latent space conditioned on codec tokens, lyrics, and text captions. For cover song generation, the melody module extracts and discretizes melody cues from reference audio to guide generation while preserving the original melodic content. Finally, we investigate DPO, GRPO, and OPD as reward-based post-training strategies for hybird-LM and apply flow-based GRPO to FullDiT to improve musicality and rendering quality. Experimental results on a multilingual automatic benchmark, complemented by the Artificial Analysis Music with Vocals leaderboard, show that the proposed framework achieves competitive performance in the evaluated settings.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.