2606.17006v1 Jun 15, 2026 cs.SD

TuneJury: 음악 생성 모델의 선호도 일치 개선을 위한 개방형 평가 지표

TuneJury: An Open Metric for Improving Music Generation Preference Alignment

Yuki Mitsufuji
Yuki Mitsufuji
Citations: 1,716
h-index: 16
Yi Ma
Yi Ma
Citations: 956
h-index: 12
Haiwen Xia
Haiwen Xia
Citations: 33
h-index: 2
Chris Donahue
Chris Donahue
Citations: 267
h-index: 4
Junghyun Koo
Junghyun Koo
Sony AI / Sony Research
Citations: 350
h-index: 9
Yonghyun Kim
Yonghyun Kim
Georgia Institute of Technology
Citations: 11
h-index: 2
Junwon Lee
Junwon Lee
Citations: 42
h-index: 3
Koichi Saito
Koichi Saito
Citations: 81
h-index: 3

본 논문에서는 텍스트-음악 변환 모델에서 텍스트 프롬프트와 오디오 클립으로부터 음악 선호도를 예측하는, 개방형 인스턴스 레벨 쌍대 보상 모델인 TuneJury를 소개합니다. 공개된 체크포인트는 아레나 스타일(A vs. B) 투표, 메트릭 일치 선호도 쌍, 크라우드 소싱 기반 쌍대 비교 및 전문가의 심미적 평가 데이터를 활용하여 학습되었습니다. 두 클립 간 예측 점수 차이는 별도의 테스트 데이터셋에서 잘 보정되었으며, 간단한 점수 임계값을 사용하여 데이터 필터링을 지원합니다. TuneJury는 별도의 테스트 데이터셋과 외부 데이터 벤치마크 모두에서 기존 모델들과 경쟁력 있는 성능을 보여줍니다. 학습 이후 공개된 생성 모델의 경우, 각 시스템별로 사후적으로 Bradley-Terry 보정을 수행하는 '앵커 교정' 기법을 도입하여 처음부터 다시 학습하는 것보다 훨씬 효율적으로 성능 향상을 이끌어낼 수 있습니다. 동일한 고정된 보상 신호는 추론 시간 내 최적의 N개 선택, DITTO 스타일 잠재 공간 최적화 및 전문가 기반 추가 훈련 등 세 가지 후속 애플리케이션에서 일관된 성능 향상을 가져옵니다. TuneJury는 https://github.com/yonghyunk1m/TuneJury 에서 사용할 수 있습니다.

Original Abstract

We introduce TuneJury, an open, instance-level pairwise reward model for text-to-music that predicts a music preference score from a text prompt and an audio clip. The released checkpoint is trained on publicly available human-preference labels covering arena-style (A vs. B) votes, metric-alignment preference pairs, crowdsourced pairwise comparisons, and expert aesthetic ratings. The predicted score margin between two clips is well calibrated on our held-out test split, supporting data filtering via a simple score threshold. TuneJury generalizes to both held-out test pairs and out-of-distribution benchmarks, remaining competitive with prior baselines on the latter. For generators released after training, we introduce anchor calibration, a post-hoc, per-system Bradley-Terry calibration that recovers agreement at substantially better data efficiency than from-scratch retraining. The same frozen reward drives consistent reward-axis gains across three downstream applications: inference-time best-of-N selection, DITTO-style latent optimization, and expert-iteration post-training. TuneJury is available at https://github.com/yonghyunk1m/TuneJury.

1 Citations
0 Influential
41.540251005511 Altmetric
6.9 Score
Original PDF
14

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!