2607.26465v1 Jul 29, 2026 cs.AI

MultivationBench: 다중 모드 순차적 동기 추론을 위한 벤치마크

MultivationBench: A Benchmark for Multimodal Sequential Motivation Reasoning

Yifan Gao
Yifan Gao
Citations: 346
h-index: 11
Xi Yang
Xi Yang
Citations: 91
h-index: 3
Tianshi ZHENG
Tianshi ZHENG
HKUST
Citations: 392
h-index: 11
Haochen Shi
Haochen Shi
Citations: 214
h-index: 8
Ka-po. Chung
Ka-po. Chung
Citations: 12
h-index: 2
Chunkit Chan
Chunkit Chan
Hong Kong University of Science and Technology
Citations: 736
h-index: 15
Yauwai Yim
Yauwai Yim
Citations: 172
h-index: 6
Qing Zong
Qing Zong
Citations: 151
h-index: 6
Kai-Huang Wong
Kai-Huang Wong
Citations: 0
h-index: 0
J. Hsiao
J. Hsiao
Citations: 23
h-index: 3
Yuxuan Liu
Yuxuan Liu
Citations: 0
h-index: 0
Weiqi Wang
Weiqi Wang
Tencent
Citations: 1,014
h-index: 18
Yixuan Fu
Yixuan Fu
Citations: 0
h-index: 0
Haoqin Liang
Haoqin Liang
Citations: 0
h-index: 0
Yangqiu Song
Yangqiu Song
Citations: 738
h-index: 15

다중 모드 대규모 언어 모델은 사회적 지능 잠재력으로 인해 큰 관심을 받고 있지만, 이러한 모델의 순차적 동기 추론 능력은 충분히 연구되지 않았습니다. 기존 평가 방법은 주로 정적인 텍스트 또는 개별적인 시각 정보만을 활용하며, 이는 실제 행동의 근본적인 원인을 종합적으로 반영하지 못합니다. 이러한 문제점을 해결하기 위해, 우리는 이야기 기반 시각적 내러티브에서 다중 모드 동기 추론을 엄격하게 평가하도록 설계된 벤치마크인 MultivationBench를 소개합니다. 이 벤치마크는 Maslow의 욕구 계층 및 Reiss의 기본 욕구와 같은 기존 심리학 이론에 기반하며, 모델이 누적된 다중 모드 정보를 통합하여 변화하는 동기를 추론하도록 요구합니다. 결과는 MultivationBench가 상당한 난제를 제시한다는 것을 보여줍니다. 테스트된 모든 모델은 순차적인 맥락에서 일관된 동기 추론을 유지하는 데 어려움을 겪으며, 이는 정적인 인식 능력과 인간다운 사회적 이해에 필수적인 역동적인 추론 간의 중요한 차이를 드러냅니다.

Original Abstract

Multimodal Large Language Models have sparked significant interest due to their potential for social intelligence; however, their ability to perform sequential motivation reasoning remains insufficiently studied. Existing evaluations predominantly examine static text or isolated visual snapshots, which do not reflect the cumulative nature of real-world behavioral drivers. To address this gap, we introduce MultivationBench, a benchmark designed to rigorously evaluate multimodal motivation reasoning within story-driven visual narratives. The benchmark builds upon established psychological frameworks - Maslow's hierarchy and Reiss's basic desires - and requires models to integrate accumulated multimodal context to infer evolving motivations. Results indicate that MultivationBench presents a significant challenge: all tested models struggle to maintain consistent motivation reasoning across sequential contexts, revealing a critical disconnect between static recognition capabilities and the dynamic reasoning essential for human-like social understanding.

0 Citations
0 Influential
9 Altmetric
45.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!