2608.05817v1 Aug 06, 2026 cs.CL

M³R-Bench: 증거 기반 다중 모드 은유 이해를 위한 통합 벤치마크

M$^3$R-Bench: A Unified Benchmark for Evidence-Grounded Multimodal Metaphor Understanding

Yuming Yang
Yuming Yang
Citations: 13
h-index: 2
Kaiwen Wei
Kaiwen Wei
Citations: 64
h-index: 5
Xiao Sun
Xiao Sun
Citations: 58
h-index: 2
Junnan Zhu
Junnan Zhu
Citations: 12
h-index: 2
Nayu Liu
Nayu Liu
Citations: 38
h-index: 3
Jingwang Huang
Jingwang Huang
Citations: 16
h-index: 2
Hao Wu
Hao Wu
Citations: 26
h-index: 2
Hongye Jiang
Hongye Jiang
Citations: 0
h-index: 0
Jiang Zhong
Jiang Zhong
Citations: 9
h-index: 1
Ruirui Chen
Ruirui Chen
Citations: 80
h-index: 5
Xinyi Jiang
Xinyi Jiang
Citations: 0
h-index: 0
Jing Shi
Jing Shi
Citations: 0
h-index: 0

은유는 감정적 태도를 전달하면서 동시에 서로 다른 영역 간의 연결을 통해 추상적인 개념에 대한 이해를 가능하게 합니다. 다중 모드 환경에서 시각 및 텍스트 정보는 목표-원천 매핑을 함께 구성하며, 이는 개념적 이해와 초모달 추론을 모두 필요로 합니다. 그러나 기존 벤치마크는 주로 독립적인 하위 작업을 통해 은유 이해를 평가하며 증거 기반 설명을 제공하지 않아 모델이 시각 및 텍스트 단서에 근거한 매핑을 구축하는지 여부를 평가하기 어렵습니다. 이러한 한계를 극복하기 위해, 우리는 인간이 검증한 주석이 포함된 1,000개의 이미지-텍스트 인스턴스로 구성된 통합적이고 증거 기반 벤치마크인 M³R-Bench를 소개합니다. 개념 은유 이론 및 비문자 언어 이해 이론에 따라, M³R-Bench는 은유 발생, 목표-원천 매핑, 감성 및 "증거 식별 - 매핑 구축 - 감성 추론" 단계를 따르는 단계별 설명을 위한 통합 주석을 제공합니다. M³R-Bench에 대한 평가는 기존 모델이 종종 시각적 증거를 간과하고 표면적인 텍스트 단서에 의존하며 부정확한 목표-원천 매핑을 생성한다는 것을 보여주며, 이는 초모달 증거-매핑 불일치를 드러냅니다. 이러한 불일치를 해결하기 위해, 우리는 과제별 강화 학습과 함께 커리큘럼 기반 추론 감독을 결합하여 모델의 추론을 은유 해석에 맞추는 M³R-Reasoner를 제안합니다. 실험 결과, 80억 개의 파라미터를 가진 비교적 작은 모델인 M³R-Reasoner가 네 가지 통합 작업 지표에서 더 큰 독점 다중 모드 언어 모델(MLLM)보다 뛰어난 성능을 보였으며, GPT-5.5에 비해 시각적 증거 및 감성 정당화 점수를 각각 28.45점과 30.11점 향상시켰습니다. 또한 Claude-Sonnet-4.6보다 평균 Rubric 점수에서 8.00점을 더 높은 성능을 보였습니다. 데이터셋 및 코드는 https://github.com/hongshi4/M3R-Bench 에서 확인할 수 있습니다.

Original Abstract

Metaphor enables the understanding of abstract concepts through cross-domain mappings while conveying affective attitudes. In multimodal scenarios, visual and textual information jointly construct Target--Source mappings, requiring both conceptual understanding and cross-modal reasoning. However, existing benchmarks mainly evaluate metaphor understanding through isolated subtasks and lack evidence-grounded explanations, making it difficult to assess whether models establish mappings grounded in visual and textual cues.To address these limitations, we introduce M$^3$R-Bench, a unified and evidence-grounded benchmark containing 1,000 image--text instances with human-verified annotations. Guided by Conceptual Metaphor Theory and theories of nonliteral language understanding, M$^3$R-Bench provides joint annotations for metaphor occurrence, Target--Source mapping, sentiment, and stage-wise explanations following ``evidence identification--mapping establishment--sentiment inference.''Evaluations on M$^3$R-Bench reveal that existing models often overlook visual evidence, rely on superficial textual cues, and produce inaccurate Target--Source mappings, exposing a cross-modal evidence--mapping mismatch. To address this mismatch, we propose M$^3$R-Reasoner, which combines curriculum-based reasoning supervision with task-aware reinforcement learning to align model reasoning with metaphor interpretation. Experiments show that, with only an 8B-parameter backbone, M$^3$R-Reasoner outperforms larger proprietary MLLMs across four unified-task metrics and improves Visual Evidence and Sentiment Justification scores over GPT-5.5 by 28.45 and 30.11 points, respectively, while surpassing Claude-Sonnet-4.6 by 8.00 points in mean rubric score. The dataset and code are available at https://github.com/hongshi4/M3R-Bench.

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!