2607.27895v1 Jul 30, 2026 cs.AI

MMHBench: 장편 동영상 내 정신 건강 이해를 위한 다각적 벤치마크

MMHBench: A Multi-Perspective Benchmark for Mental Health Understanding in Long-Form Videos

Jinpeng Hu
Jinpeng Hu
Hefei University of Technology
Citations: 797
h-index: 15
Peipei Song
Peipei Song
Citations: 55
h-index: 4
Erqiang Wang
Erqiang Wang
Citations: 0
h-index: 0
Shan A. Wang
Shan A. Wang
Citations: 0
h-index: 0
Zhuo Li
Zhuo Li
Citations: 39
h-index: 3
Xun Yang
Xun Yang
Citations: 17
h-index: 1
Meng Wang
Meng Wang
Citations: 33
h-index: 2

장편 동영상에서 정신 건강을 이해하는 것은 관찰 가능한 행동, 대인 관계 맥락 및 잠재적인 심리 상태에 대한 미묘한 추론이 필요합니다. 기존의 벤치마크는 이러한 작업을 대부분 단순화된 분류로 줄여 모델이 실제로 심리 현상을 이해하는지, 아니면 피상적인 상관관계에 의존하는지에 대한 제한적인 통찰력을 제공합니다. 이러한 한계를 해결하기 위해, 우리는 268개의 장편 동영상과 2,184개의 신중하게 구성된 질문으로 구성된 다각적 정신 건강 이해를 위한 종합적인 멀티모달 벤치마크인 MMHBench를 소개합니다. MMHBench는 평가를 두 가지 상호 보완적인 방식으로 구성합니다: (1) 제3자 관점 평가, 이는 관찰 가능한 행동 및 멀티모달 증거 해석에 중점을 둔 605개의 질문으로 구성되며, (2) 1인칭 관점 수용, 이는 가용한 멀티모달 증거에 의해 뒷받침되는 정신 상태의 해석을 파악하기 위해 관점에 따른 추론이 필요한 1,579개의 질문으로 구성됩니다. 우리는 다양한 사회적 역할을 시뮬레이션하여 여러 관점에서 질문을 생성하는 Multi-Agent Question Generation (MAQG) 프레임워크를 제안합니다. 생성된 질문은 다중 역할 피드백 및 반복적인 최적화를 통해 개선되고, 품질과 유효성을 보장하기 위해 전문가의 검토를 거칩니다. 오픈 소스 모델과 선도적인 비공개 모델을 포함한 22개의 대표적인 멀티모달 대규모 언어 모델(MLLM)에 대한 광범위한 평가는 장편 동영상 정신 건강 이해가 여전히 매우 어려운 과제임을 보여줍니다.

Original Abstract

Mental health understanding in long-form videos requires nuanced reasoning over observable behavior, interpersonal context, and latent psychological states. Existing benchmarks largely reduce this task to coarse-grained classification, providing limited insight into whether models truly understand psychological phenomena or rely on superficial correlations. To address this limitation, we introduce MMHBench, a comprehensive multimodal benchmark for multi-perspective mental health understanding, comprising 268 long-form videos and 2,184 carefully curated questions. MMHBench organizes the evaluation into two complementary settings: (1) third-person assessment, consisting of 605 questions that focus on the interpretation of observable behaviors and multimodal evidence, and (2) first-person perspective-taking, comprising 1,579 questions that require perspective-conditioned reasoning to identify the interpretation of the mental state supported by the available multimodal evidence. We propose a Multi-Agent Question Generation (MAQG) framework that simulates diverse social roles to synthesize questions from multiple perspectives. The generated questions are refined through multi-role feedback and iterative optimization, followed by expert-guided verification to ensure quality and validity. Extensive evaluation of 22 representative multimodal large language models (MLLMs), spanning both open-source and leading closed-source models, demonstrates that long-form video mental health understanding remains highly challenging.

0 Citations
0 Influential
7.5 Altmetric
37.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!