2601.19673v1 Jan 27, 2026 cs.SD

다중 모드 대규모 언어 모델의 오디오 추론 능력 평가를 위한 벤치마크

A Benchmark for Audio Reasoning Capabilities of Multimodal Large Language Models

Iwona Christop
Iwona Christop
Citations: 16
h-index: 2
Mateusz Czyznikiewicz
Mateusz Czyznikiewicz
Citations: 3
h-index: 1
Pawel Sk'orzewski
Pawel Sk'orzewski
Citations: 2
h-index: 1
Lukasz Bondaruk
Lukasz Bondaruk
Citations: 6
h-index: 2
Jakub Kubiak
Jakub Kubiak
Citations: 6
h-index: 2
Marcin Lewandowski
Marcin Lewandowski
Citations: 2
h-index: 1
Marek Kubis
Marek Kubis
Citations: 23
h-index: 3

현재 다중 모드 대규모 언어 모델의 오디오 모드를 테스트하는 벤치마크는 일반적으로 화자 식별 또는 성별 판별과 같은 다양한 오디오 작업을 개별적으로 테스트하는 데 집중합니다. 이러한 벤치마크는 다중 모드 모델이 다양한 범주의 오디오 작업을 결합하여 추론 능력을 필요로 하는 질문에 답변할 수 있는지 여부를 검증하는 데 한계가 있습니다. 이러한 문제를 해결하기 위해, 우리는 오디오 신호에 대한 추론 능력을 평가하기 위한 새로운 벤치마크인 오디오 추론 작업(Audio Reasoning Tasks, ART)을 제안합니다.

Original Abstract

The present benchmarks for testing the audio modality of multimodal large language models concentrate on testing various audio tasks such as speaker diarization or gender identification in isolation. Whether a multimodal model can answer the questions that require reasoning skills to combine audio tasks of different categories, cannot be verified with their use. To address this issue, we propose Audio Reasoning Tasks (ART), a new benchmark for assessing the ability of multimodal models to solve problems that require reasoning over audio signal.

1 Citations
0 Influential
1.5 Altmetric
8.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!