2601.21666v1 Jan 29, 2026 cs.AI

SONIC-O1: 오디오-비디오 이해에 대한 멀티모달 대규모 언어 모델 평가를 위한 실세계 벤치마크

SONIC-O1: A Real-World Benchmark for Evaluating Multimodal Large Language Models on Audio-Video Understanding

Ahmed Y. Radwan
Ahmed Y. Radwan
Citations: 43
h-index: 2
Christos Emmanouilidis
Christos Emmanouilidis
Citations: 96
h-index: 4
Hina Tabassum
Hina Tabassum
Citations: 75
h-index: 5
D. Pandya
D. Pandya
Citations: 247
h-index: 10
Shaina Raza
Shaina Raza
Citations: 23
h-index: 3

멀티모달 대규모 언어 모델(MLLM)은 최근 AI 연구의 주요 초점입니다. 그러나 대부분의 선행 연구는 정적 이미지 이해에 집중되어 있으며, 순차적인 오디오-비디오 데이터를 처리하는 능력은 여전히 충분히 탐구되지 않았습니다. 이러한 공백은 실세계 환경에서 MLLM의 성능을 체계적으로 평가하기 위한 고품질 벤치마크의 필요성을 부각시킵니다. 우리는 13개의 실세계 대화 도메인에 걸쳐 4,958개의 주석과 인구통계학적 메타데이터를 포함하는, 포괄적이고 완전히 사람이 검증한 벤치마크인 SONIC-O1을 소개합니다. SONIC-O1은 개방형 요약, 객관식 질문(MCQ) 답변, 근거(추론)를 포함한 시간적 위치 파악(temporal localization) 등 주요 작업에 대해 MLLM을 평가합니다. 폐쇄형(closed-source) 및 오픈 소스 모델에 대한 실험을 통해 몇 가지 한계점이 드러났습니다. 두 모델군 간의 MCQ 정확도 성능 격차는 상대적으로 작았지만, 최고 성능의 폐쇄형 모델과 오픈 소스 모델 사이에는 시간적 위치 파악 작업에서 22.6%의 상당한 성능 차이가 관찰되었습니다. 또한 인구통계학적 그룹 전반에 걸쳐 성능이 추가로 저하되는 현상은 모델 동작에 지속적인 불균형이 있음을 나타냅니다. 종합적으로 SONIC-O1은 시간적 근거가 명확하고 사회적으로 견고한 멀티모달 이해를 위한 개방형 평가 제품군을 제공합니다. 우리는 재현성과 연구를 위해 SONIC-O1을 공개합니다: 프로젝트 페이지: https://vectorinstitute.github.io/sonic-o1/ 데이터셋: https://huggingface.co/datasets/vector-institute/sonic-o1 Github: https://github.com/vectorinstitute/sonic-o1 리더보드: https://huggingface.co/spaces/vector-institute/sonic-o1-leaderboard

Original Abstract

Multimodal Large Language Models (MLLMs) are a major focus of recent AI research. However, most prior work focuses on static image understanding, while their ability to process sequential audio-video data remains underexplored. This gap highlights the need for a high-quality benchmark to systematically evaluate MLLM performance in a real-world setting. We introduce SONIC-O1, a comprehensive, fully human-verified benchmark spanning 13 real-world conversational domains with 4,958 annotations and demographic metadata. SONIC-O1 evaluates MLLMs on key tasks, including open-ended summarization, multiple-choice question (MCQ) answering, and temporal localization with supporting rationales (reasoning). Experiments on closed- and open-source models reveal limitations. While the performance gap in MCQ accuracy between two model families is relatively small, we observe a substantial 22.6% performance difference in temporal localization between the best performing closed-source and open-source models. Performance further degrades across demographic groups, indicating persistent disparities in model behavior. Overall, SONIC-O1 provides an open evaluation suite for temporally grounded and socially robust multimodal understanding. We release SONIC-O1 for reproducibility and research: Project page: https://vectorinstitute.github.io/sonic-o1/ Dataset: https://huggingface.co/datasets/vector-institute/sonic-o1 Github: https://github.com/vectorinstitute/sonic-o1 Leaderboard: https://huggingface.co/spaces/vector-institute/sonic-o1-leaderboard

3 Citations
0 Influential
33.047189562171 Altmetric
13.9 Score
Original PDF
4

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!