InvFlowFD: 플로우 매칭 역전산을 이용한 참조 데이터 및 배경 집합이 없는 음악 품질 평가 지표
InvFlowFD: Reference-Free and Background-Set-Free Perceptual Music Quality Metric with Flow Matching Inversion
기존의 참조 데이터가 필요 없는 음악 품질 평가 방법들은 노이즈-깨끗 오디오 쌍 데이터를 사용할 필요성을 줄여주지만, 여전히 깨끗한 오디오 샘플들의 통계 정보를 계산하는 데 사용되는 배경 집합에 의존합니다. 본 연구에서는 사전 훈련된 플로우 매칭 모델만을 사용하여 배경 집합 없이 참조 데이터가 없는 품질 평가 방법을 제안합니다. 간단한 Euler 적분을 통한 조건부 플로우 매칭 역전산이 다양한 인공적인 왜곡을 감지하고, 음악 생성 모델의 품질을 인간의 청각 판단과 정확하게 비교하는 데 충분하다는 것을 보여줍니다. 우리는 InvFlowFD를 소개하며, 이는 플로우 역전산을 수행하고, 역전된 샘플 그룹을 사전 분포와 비교합니다. 본 연구에서는 제안하는 방법을 기존 방법들과 정량적으로 비교하고, 심층적인 인간 평가 연구를 통해 그 성능을 검증했습니다. 결과는 InvFlowFD가 소리 왜곡의 인간 인지 및 생성 모델의 품질과 높은 상관관계를 가지며, 기존 지표보다 더 유연하고 제약이 적다는 것을 시사합니다.
Existing reference-free methods for evaluating music perceptual quality alleviate the need for paired noisy-clean data, but they still rely on a background set, which is used to compute aggregated statistics of clean audio samples. In this work, we propose a novel approach that eliminates this requirement, achieving background-set-free and reference-free quality estimation using only a pre-trained Flow Matching backbone. We demonstrate that unconditional Flow Matching inversion via simple Euler integration is sufficient to detect various artificial distortions and accurately rank music generation models against human perceptual judgments. We introduce InvFlowFD, which performs flow inversion and compares a group of inverted samples to the prior distribution. We evaluate our method against prior work, quantitatively and with a thorough human study. Results suggest that InvFlowFD is highly correlated with human perception of sound distortions, as well as generative models' quality, while being more flexible and less restrictive than existing metrics.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.