TuneJury: 음악 생성 모델의 선호도 일치 개선을 위한 개방형 평가 지표
TuneJury: An Open Metric for Improving Music Generation Preference Alignment
본 논문에서는 텍스트-음악 변환 모델에서 텍스트 프롬프트와 오디오 클립으로부터 음악 선호도를 예측하는, 개방형 인스턴스 레벨 쌍대 보상 모델인 TuneJury를 소개합니다. 공개된 체크포인트는 아레나 스타일(A vs. B) 투표, 메트릭 일치 선호도 쌍, 크라우드 소싱 기반 쌍대 비교 및 전문가의 심미적 평가 데이터를 활용하여 학습되었습니다. 두 클립 간 예측 점수 차이는 별도의 테스트 데이터셋에서 잘 보정되었으며, 간단한 점수 임계값을 사용하여 데이터 필터링을 지원합니다. TuneJury는 별도의 테스트 데이터셋과 외부 데이터 벤치마크 모두에서 기존 모델들과 경쟁력 있는 성능을 보여줍니다. 학습 이후 공개된 생성 모델의 경우, 각 시스템별로 사후적으로 Bradley-Terry 보정을 수행하는 '앵커 교정' 기법을 도입하여 처음부터 다시 학습하는 것보다 훨씬 효율적으로 성능 향상을 이끌어낼 수 있습니다. 동일한 고정된 보상 신호는 추론 시간 내 최적의 N개 선택, DITTO 스타일 잠재 공간 최적화 및 전문가 기반 추가 훈련 등 세 가지 후속 애플리케이션에서 일관된 성능 향상을 가져옵니다. TuneJury는 https://github.com/yonghyunk1m/TuneJury 에서 사용할 수 있습니다.
We introduce TuneJury, an open, instance-level pairwise reward model for text-to-music that predicts a music preference score from a text prompt and an audio clip. The released checkpoint is trained on publicly available human-preference labels covering arena-style (A vs. B) votes, metric-alignment preference pairs, crowdsourced pairwise comparisons, and expert aesthetic ratings. The predicted score margin between two clips is well calibrated on our held-out test split, supporting data filtering via a simple score threshold. TuneJury generalizes to both held-out test pairs and out-of-distribution benchmarks, remaining competitive with prior baselines on the latter. For generators released after training, we introduce anchor calibration, a post-hoc, per-system Bradley-Terry calibration that recovers agreement at substantially better data efficiency than from-scratch retraining. The same frozen reward drives consistent reward-axis gains across three downstream applications: inference-time best-of-N selection, DITTO-style latent optimization, and expert-iteration post-training. TuneJury is available at https://github.com/yonghyunk1m/TuneJury.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.