MuseBench: 멀티모달 대규모 언어 모델(MLLM)의 의도 기반 시청각 예술 이해 성능 평가
MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs
시청각 예술은 영화, 미술, 연극, 게임 디자인 등 다양한 창작 분야를 포괄하며, 공포감을 조성하기 위해 좁은 화면을 사용하거나, 슬픔을 표현하기 위해 침묵과 클로즈업 장면을 활용하는 것처럼, 시각적, 청각적 요소와 스토리텔링이 결합되어 예술적인 의미가 형성됩니다. 진정한 예술 이해는 단순히 묘사된 내용을 인지하는 것을 넘어, 특정 창작 선택을 통해 왜 그러한 표현이 사용되었는지 추론하는 능력을 포함합니다. 최근 멀티모달 대규모 언어 모델(MLLM)의 발전에도 불구하고, 이러한 예술적 이해의 중요한 측면은 기존 평가 지표에서 충분히 다루어지지 않고 있습니다. 대부분의 기존 평가 지표는 시각적인 인지 능력만을 측정하며, 창작 의도에 대한 추론 능력은 간과하는 경향이 있습니다. 이러한 문제점을 해결하기 위해, 본 연구에서는 MLLM의 미묘한 예술적 이해 능력을 평가할 수 있는 종합적인 벤치마크인 MuseBench를 제안합니다. MuseBench는 영화, 정지 영상 미술, 연극, 게임 아트 분야에 걸쳐 총 4,016개의 질문으로 구성되어 있으며, 전문가 해설과 시각 자료를 결합한 1만 개 이상의 비디오 에세이에서 추출되었습니다. 예술 분석의 개방형 특성을 반영하기 위해, 객관식 문제와 함께 다양한 선택지를 제공하는 객관식 다중선택 문제 형식을 사용했습니다. 모든 질문은 단답형 필터링, 적대적 오답 생성, 전문가 검증을 포함한 4단계 반복적인 과정을 거쳐 생성 및 개선되었습니다. 최첨단 MLLM 28개를 대상으로 실시한 초기 평가 결과, 가장 높은 성능을 보이는 모델도 48.29%의 정확도를 기록했으며, 이는 인간 전문가의 평균 87.18%에 비해 현저히 낮은 수치입니다. 이러한 결과는 현재 모델들이 예술 분야에 대한 전문성이 부족하다는 것을 보여줍니다.
Audiovisual arts encompass diverse creative disciplines, including cinema, visual arts, stage performance, and game design, where artistic meaning arises from deliberate combinations of visual, auditory, and narrative elements (e.g., fear amplified through claustrophobic framing, or grief conveyed through silence and lingering close-ups). True artistic understanding extends beyond recognizing what is depicted to reasoning about why it is expressed through particular creative choices. Despite the strong progress of multimodal large language models (MLLMs), this critical aspect of artistic understanding remains underexplored, as existing benchmarks largely measure perceptual recognition while overlooking reasoning about creative intent. To address this gap, we introduce Musebench, a comprehensive benchmark designed to evaluate MLLMs on nuanced artistic understanding. It comprises 4,016 questions spanning cinematic arts, static visual arts, stage performing arts, and game arts, distilled from over 10K candidate video essays that pair professional commentary with visual demonstration. To capture the open-ended nature of artistic analysis at scale, the benchmark combines single-select and variable-option multi-select questions. All questions are generated and refined through a four-phase iterative pipeline combining shortcut filtering, adversarial distractors, and expert validation. Comprehensive zero-shot evaluation of 28 state-of-the-art MLLMs reveals that even the best-performing model achieves only 48.29% accuracy, substantially below human expert performance of 87.18%, exposing a significant gap in current models' creative domain expertise.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.