대규모 오디오 언어 모델의 사실 기반 음악 이해도 평가
Assessing Factual Music Comprehension in Large Audio Language Models
대규모 오디오 언어 모델(LALM)은 다중 모드 표현을 활용하여 오디오에 대한 자연어 질문에 대해 개방형 답변을 생성합니다. 본 논문에서는 (1) 인기 있는 MusicQA 데이터 세트를 사용하여 LALM을 평가하는 것이 모델의 음악 관련 응답이 사실적으로 정확한지 측정하지 못한다는 경험적 증거를 제시하고, (2) LALM의 음악 이해 능력을 평가하기 위한 새로운 프로토콜을 개발합니다. 구체적으로, 우리는 LALM에게 사실적으로 검증 가능한 정보를 요청하고, LALM의 개방형 응답을 정밀도, 재현율 및 F1 점수를 사용하여 객관적으로 평가할 수 있는 구조화된 형식으로 파싱하는 평가 프로토콜을 제안합니다. 이 프로토콜을 사용하여 MusicNet, Free Music Archive 및 OverClocked ReMix라는 세 가지 다양한 데이터 세트에 정의된 6가지 사실 정보 검색 작업으로 구성된 벤치마크를 정의했습니다. 우리는 Gemini와 같은 최첨단 모델과 Music Flamingo와 같은 최신 오픈 소스 모델을 포함한 9개의 최근 LALM을 벤치마킹하고, 새로운 LALM의 벤치마킹을 용이하게 하기 위해 평가 스크립트 모음을 https://github.com/DCL2004/LALM-Eval 에서 공개합니다.
Large audio language models (LALMs) leverage multimodal representations to generate open-ended answers to natural language queries about audio. In this paper, we (1) provide empirical evidence that assessment of LALMs using the popular MusicQA dataset fails to measure whether a model's responses about music are factually correct, and (2) develop a new protocol for assessing the music comprehension capabilities of LALMs. Specifically, we propose an evaluation protocol that prompts a LALM for factually verifiable information, and parses its open-ended response into a structured format that can be objectively assessed using Precision, Recall, and F1 scores. Using this protocol, we define a benchmark consisting of six factual information retrieval tasks defined on three diverse datasets: MusicNet, the Free Music Archive, and OverClocked ReMix. We benchmark nine recent LALMs, including frontier models like Gemini and the latest open models like Music Flamingo, and release the suite of evaluation scripts at https://github.com/DCL2004/LALM-Eval to facilitate benchmarking of new LALMs.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.