2605.03485v1 May 05, 2026 cs.CV

MHPR: 대규모 비전-언어 모델을 위한 다차원 인간 이해 및 추론 벤치마크

MHPR: Multidimensional Human Perception and Reasoning Benchmark for Large Vision-Languate Models

Kangkang Wang
Kangkang Wang
Citations: 352
h-index: 3
Qinting Jiang
Qinting Jiang
Citations: 11
h-index: 2
Bo Ren
Bo Ren
Citations: 42
h-index: 1
Shengzhao Wen
Shengzhao Wen
Citations: 90
h-index: 3
Wanping Zhang
Wanping Zhang
Citations: 8
h-index: 1

다차원적인 인간 이해는 영화 분석 및 가상 디지털 휴먼과 같은 실제 응용 분야에서 필수적이지만, 현재의 대규모 비전-언어 모델 벤치마크는 대부분 단일 작업에 초점을 맞추고 있으며, 세분화된 인간 중심적인 평가가 부족합니다. 본 연구에서는 인간 중심적인 장면을 포괄적으로 평가하기 위한 벤치마크인 MHPR을 소개합니다. MHPR은 개인, 다인, 인간-객체 상호작용의 다양한 측면을 포함하며, 다단계 데이터 설계(캡션이 포함된 원시 데이터(C-RD), 지도 학습 데이터(SFT-D), 강화 학습 데이터(RL-D), 테스트 데이터(T-D))와 함께, 고품질의 확장 가능한 어노테이션을 보장하기 위한 자동 캡션/VQA 생성 파이프라인(ACVG)을 제공합니다. ACVG는 범주별 속성 분해, 속성별 재작성 및 다중 모델 투표를 수행합니다. 우리는 최첨단 비전-언어 모델을 사용하여 세분화된 속성(외모, 의상, 자세, 부분) 및 고수준 의미(사회적 관계, 동작 의미, 공간 관계, 의도 및 기능)를 평가했습니다. 우리의 연구 결과는 다음과 같습니다. 1) 형식에 맞춘 SFT 데이터는 명령어 이해 및 안정성을 크게 향상시킵니다. 2) 오답 분석에서 파생된 문제 중심적인 RL 데이터는 어려운 사례에 대한 인식 및 추론 능력을 더욱 향상시킵니다. 3) MHPR을 사용하여 Qwen2.5-VL-7B를 학습하면 상당한 성능 향상을 얻을 수 있으며, 이는 훨씬 더 큰 모델과 거의 동등한 수준에 도달합니다. 우리는 인간 중심적인 인식 및 추론에 대한 재현 가능하고 확장 가능한 연구를 촉진하기 위해 ACVG와 MHPR을 공개합니다.

Original Abstract

Multidimensional human understanding is essential for real-world applications such as film analysis and virtual digital humans, yet current LVLM benchmarks largely focus on single-task settings and lack fine-grained, human-centric evaluation. In this work, we introduce MHPR, a comprehensive benchmark for joint perception-reasoning over human-centric scenes spanning individual, multi-person, and human-object interaction dimensions. MHPR comprises a multi-level data design-Captioned Raw Data (C-RD), Supervised Fine-Tuning Data (SFT-D), Reinforcement Learning Data (RL-D), and Test Data (T-D)-together with an automated caption/VQA generation pipeline (ACVG) that performs category-wise attribute decomposition, attribute-specific rewriting, and multi-model voting to ensure high-quality, scalable annotations. We evaluate state-of-the-art vision-language models on fine-grained attributes (appearance, clothing, pose, parts) and high-level semantics (social relations, action semantics, spatial relations, intent and functionality). Our findings show that: 1) format-aligned SFT data substantially improves instruction following and stability; 2) challenge-focused RL data derived from bad-case analysis further enhances perception and reasoning on difficult instances; and 3) training Qwen2.5-VL-7B with MHPR yields significant gains, achieving near-parity with considerably larger models. We release ACVG and MHPR to facilitate reproducible, extensible research on human-centric perception and reasoning.

0 Citations
0 Influential
1.5 Altmetric
7.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!