다중 모드 대규모 언어 모델의 오디오 추론 능력 평가를 위한 벤치마크
A Benchmark for Audio Reasoning Capabilities of Multimodal Large Language Models
현재 다중 모드 대규모 언어 모델의 오디오 모드를 테스트하는 벤치마크는 일반적으로 화자 식별 또는 성별 판별과 같은 다양한 오디오 작업을 개별적으로 테스트하는 데 집중합니다. 이러한 벤치마크는 다중 모드 모델이 다양한 범주의 오디오 작업을 결합하여 추론 능력을 필요로 하는 질문에 답변할 수 있는지 여부를 검증하는 데 한계가 있습니다. 이러한 문제를 해결하기 위해, 우리는 오디오 신호에 대한 추론 능력을 평가하기 위한 새로운 벤치마크인 오디오 추론 작업(Audio Reasoning Tasks, ART)을 제안합니다.
The present benchmarks for testing the audio modality of multimodal large language models concentrate on testing various audio tasks such as speaker diarization or gender identification in isolation. Whether a multimodal model can answer the questions that require reasoning skills to combine audio tasks of different categories, cannot be verified with their use. To address this issue, we propose Audio Reasoning Tasks (ART), a new benchmark for assessing the ability of multimodal models to solve problems that require reasoning over audio signal.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.