LeVo 2: 계층적 표현 모델링 및 점진적인 추가 학습을 통한 안정적이고 아름다운 노래 생성
LeVo 2: Stable and Melodious Song Generation via Hierarchical Representation Modeling and Progressive Post-Training
전체 길이의 노래 생성을 위해서는 일관성 유지, 음악성 구현, 상세한 보컬 및 반주 음향 렌더링, 그리고 가사와 프롬프트 준수가 필수적입니다. 기존의 언어 모델 기반 시스템은 구조적인 트레이드오프에 직면합니다: 토큰 혼합 모델링은 보컬-악기 조화를 유지하지만 트랙별 세부 정보를 흐리게 하고, 이중 트랙 예측은 음향을 개선하지만 더 긴 시퀀스가 필요하고 전반적인 계획 능력을 약화시킵니다. 본 논문에서는 제어 가능한 전체 길이의 노래 생성을 위한 하이브리드 LLM-Diffusion 프레임워크인 LeVo 2를 제시합니다. LeVo 2는 이러한 트레이드오프를 계층적 모델링으로 해결합니다. LeLM은 먼저 의미론적 계획을 위해 혼합 토큰을 예측한 다음, 각 트랙의 세부적인 개선을 위해 보컬 및 반주 토큰을 병렬로 예측하며, Diffusion 기반 Music Codec는 전체 길이의 파형을 재구성합니다. 본 논문의 핵심 기여점은 정렬을 위한 미학 지향적인 학습 스케줄입니다. 사전 학습 단계에서 자동화된 음악 미학 평가 프레임워크가 대규모 데이터에 음악성 수준 조건을 할당하여, 선호도 정렬 전에 음악적 우선순위를 제공합니다. 점진적인 추가 학습을 통해 SFT(Supervised Fine-Tuning), 대규모 오프라인 DPO(Direct Preference Optimization), 그리고 폐루프 준온라인 DPO를 적용하여 생성 품질, 제어 가능성 및 음악성을 개별적으로 개선합니다. 모듈식 확장을 통해 트랙별 LM(Language Model)을 훈련하여 음향을 개선하는 동시에 정렬된 의미론적 계획 기능을 유지합니다. 이러한 스케줄은 음악성 학습, 제어 가능성 정렬 및 음향 개선을 분리하여 최적화 충돌과 정적인 오프라인 선호도 쌍의 한계를 완화합니다. 전문가 청취 테스트와 객관적인 평가 결과는 LeVo 2가 여섯 가지 주관적인 측면에서 오픈 소스 기준 모델보다 우수한 성능을 보이며, 여러 청취 지표에서 선도적인 상용 시스템에 근접하는 성능을 보이는 것을 확인했습니다. 추가 실험을 통해 학습 전략, 미학 가이드, 확장 및 계층적 구조의 효과를 검증했습니다.
Full-length song generation must preserve coherence and musicality, render detailed vocal and accompaniment acoustics, and follow lyrics and prompts. Existing language model-based systems face a structural trade-off: mixed-token modeling preserves vocal-instrument coordination but obscures track-specific details, whereas dual-track prediction improves acoustics but requires longer sequences and weakens global planning. We present LeVo 2, a hybrid LLM-Diffusion framework for controllable full-length song generation. LeVo 2 formulates this trade-off as hierarchical modeling: LeLM first predicts mixed tokens for semantic planning, then predicts vocal and accompaniment tokens in parallel for track-specific refinement, while a diffusion-based Music Codec reconstructs full-length waveforms. A central contribution of this extended version is an aesthetics-guided training schedule for alignment. During pre-training, an automated music aesthetic evaluation framework assigns musicality-tier conditions to large-scale data, providing musicality priors before preference alignment. Progressive post-training applies SFT, large-scale offline DPO, and closed-loop semi-online DPO to separately improve generation quality, controllability, and musicality. Modular extension then trains the Track-Specific LM for acoustic refinement while preserving the aligned semantic planner. This schedule separates musicality learning, controllability alignment, and acoustic refinement, mitigating optimization conflict and the limitations of static offline preference pairs. Expert listening tests and objective evaluations show that LeVo 2 outperforms open-source baselines across six subjective dimensions, and approaches leading commercial systems on several listening metrics. Ablations validate the effects of the training strategy, aesthetics guidance, scaling, and hierarchical architecture.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.