MMAC: 오디오 캡셔닝을 위한 방대한 다차원 벤치마크
MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning
오디오 대규모 언어 모델(AudioLLMs)의 발전과 함께, 오디오 캡셔닝은 간략한 설명에서 벗어나 개방적이고 세밀한 자유 형식 설명을 지향해야 합니다. 기존 평가 방법은 주로 생성 품질이나 작업 성능에 초점을 맞추기 때문에 정보 포함 범위 및 설명 신뢰성을 진단하기 어렵습니다. 본 연구에서는 오디오 캡셔닝을 위한 방대한 다차원 벤치마크인 MMAC을 제안합니다. MMAC은 20개 이상의 데이터 소스에서 수집된 5,638개의 오디오 클립으로 구성되어 있으며, 6가지 기능 범주와 15가지 평가 차원을 포함합니다. MMAC은 모델이 생성한 설명에 대해, 대상 차원에서 관련 정보가 언급되었는지 여부와 언급된 내용이 참조 레이블과 일치하는지 여부를 확인합니다. 대표적인 오픈 소스 및 독점 AudioLLM을 평가한 결과, 평가 차원, 정보 포함 범위 및 설명 신뢰성 측면에서 뚜렷한 차이가 나타났습니다. MMAC 벤치마크와 평가 코드를 공개할 예정입니다.
With the development of audio large language models (AudioLLMs), audio captioning needs to move from brief descriptions toward open-ended and fine-grained free-form descriptions. Existing evaluations often focus on generation quality or task performance, making it difficult to diagnose information coverage and description reliability. We propose MMAC, a \textbf{M}assive \textbf{M}ulti-dimensional benchmark for \textbf{A}udio \textbf{C}aptioning. MMAC contains 5,638 audio clips from more than 20 data sources, covering 6 capability categories and 15 evaluation dimensions. Given a model-generated caption, MMAC checks whether it mentions relevant information in the target dimension and whether the mentioned content is consistent with the reference label. We evaluate representative open-source and proprietary AudioLLMs. Results show clear differences across evaluation dimensions, information coverage, and description reliability. We will release the MMAC benchmark and evaluation code.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.