ArtECulture: 다중 모달 대규모 언어 모델에서 문화적 맥락을 고려한 시각적 감정 이해의 성능 평가
ArtECulture: Benchmarking Culture-Conditioned Visual Emotion Understanding in Multimodal Large Language Models
기존의 시각적 감정 이해 방법은 일반적으로 감정 인식에서의 문화적 차이를 간과합니다. 본 연구에서는 문화적 맥락을 고려한 시각적 감정 이해라는 새로운 과제를 제안합니다. 이 과제는 주어진 이미지에 대한 문화별 특유의 감정 인식을 예측하고, 그 근거를 설명하는 것을 목표로 합니다. 기존의 관련 벤치마크가 존재하지만, 개별적인 주석의 일관성 부족으로 인해 문화 수준에서의 감정 라벨을 도출하는 데 어려움이 있으며, 문화적 다양성의 불균형도 문제점으로 지적됩니다. 따라서, 본 연구에서는 영어, 중국어, 아랍어를 포함한 세 가지 문화권에서 수집된 6,792점의 예술 작품에 대해 문화별 감정 라벨과 설명을 제공하는 ArtECulture라는 새로운 벤치마크를 제시합니다. 서구 및 비서구 콘텐츠의 균형을 맞추었습니다. 제로샷 설정 하에서 공개 및 비공개 다중 모달 대규모 언어 모델(MLLM) 16개를 평가한 결과, 이 과제가 여전히 어렵다는 것을 알 수 있었으며, 가장 성능이 좋은 모델도 50% 미만의 정확도를 보였습니다. 이러한 한계를 극복하기 위해, 컨셉 기반의 문화적 감정 지식 베이스를 활용하여 추가적인 학습 없이 MLLM에 명시적인 문화 정보를 주입하는 검색 증강 방식의 문화적 맥락을 고려한 감정 이해 프레임워크를 제안합니다. 이 프레임워크는 문화적으로 일관된 감정 예측과 근거 있는 설명 생성 능력을 향상시킵니다. 본 벤치마크와 코드는 공개될 예정입니다.
Existing visual emotion understanding methods typically ignore cultural variations in emotional perception. We introduce culture-conditioned visual emotion understanding, a task that predicts the culture-specific emotional perception of a given image and explains the underlying rationale. Although related benchmarks exist, they are limited by inconsistent individual annotations, which hinder the derivation of majority-supported culture-level emotion labels, and imbalanced cultural coverage. Thus, we present ArtECulture, a benchmark containing 6,792 artworks with culture-specific emotion labels and explanations across English, Chinese, and Arabic cultures, with balanced Western and non-Western content. Evaluations of 16 open- and closed-source Multimodal Large Language Models (MLLMs) under a zero-shot setting reveal that the task remains challenging, with the best model achieving below 50\% accuracy. To address this limitation, we introduce a retrieval-augmented culture-conditioned emotion understanding framework, which leverages a concept-based cultural emotion knowledge base to inject explicit cultural knowledge into MLLMs without additional training. The framework improves both culturally aligned emotion prediction and grounded explanation generation. Our benchmark and code will be publicly released.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.