MHPR: 대규모 비전-언어 모델을 위한 다차원 인간 이해 및 추론 벤치마크
MHPR: Multidimensional Human Perception and Reasoning Benchmark for Large Vision-Languate Models
다차원적인 인간 이해는 영화 분석 및 가상 디지털 휴먼과 같은 실제 응용 분야에서 필수적이지만, 현재의 대규모 비전-언어 모델 벤치마크는 대부분 단일 작업에 초점을 맞추고 있으며, 세분화된 인간 중심적인 평가가 부족합니다. 본 연구에서는 인간 중심적인 장면을 포괄적으로 평가하기 위한 벤치마크인 MHPR을 소개합니다. MHPR은 개인, 다인, 인간-객체 상호작용의 다양한 측면을 포함하며, 다단계 데이터 설계(캡션이 포함된 원시 데이터(C-RD), 지도 학습 데이터(SFT-D), 강화 학습 데이터(RL-D), 테스트 데이터(T-D))와 함께, 고품질의 확장 가능한 어노테이션을 보장하기 위한 자동 캡션/VQA 생성 파이프라인(ACVG)을 제공합니다. ACVG는 범주별 속성 분해, 속성별 재작성 및 다중 모델 투표를 수행합니다. 우리는 최첨단 비전-언어 모델을 사용하여 세분화된 속성(외모, 의상, 자세, 부분) 및 고수준 의미(사회적 관계, 동작 의미, 공간 관계, 의도 및 기능)를 평가했습니다. 우리의 연구 결과는 다음과 같습니다. 1) 형식에 맞춘 SFT 데이터는 명령어 이해 및 안정성을 크게 향상시킵니다. 2) 오답 분석에서 파생된 문제 중심적인 RL 데이터는 어려운 사례에 대한 인식 및 추론 능력을 더욱 향상시킵니다. 3) MHPR을 사용하여 Qwen2.5-VL-7B를 학습하면 상당한 성능 향상을 얻을 수 있으며, 이는 훨씬 더 큰 모델과 거의 동등한 수준에 도달합니다. 우리는 인간 중심적인 인식 및 추론에 대한 재현 가능하고 확장 가능한 연구를 촉진하기 위해 ACVG와 MHPR을 공개합니다.
Multidimensional human understanding is essential for real-world applications such as film analysis and virtual digital humans, yet current LVLM benchmarks largely focus on single-task settings and lack fine-grained, human-centric evaluation. In this work, we introduce MHPR, a comprehensive benchmark for joint perception-reasoning over human-centric scenes spanning individual, multi-person, and human-object interaction dimensions. MHPR comprises a multi-level data design-Captioned Raw Data (C-RD), Supervised Fine-Tuning Data (SFT-D), Reinforcement Learning Data (RL-D), and Test Data (T-D)-together with an automated caption/VQA generation pipeline (ACVG) that performs category-wise attribute decomposition, attribute-specific rewriting, and multi-model voting to ensure high-quality, scalable annotations. We evaluate state-of-the-art vision-language models on fine-grained attributes (appearance, clothing, pose, parts) and high-level semantics (social relations, action semantics, spatial relations, intent and functionality). Our findings show that: 1) format-aligned SFT data substantially improves instruction following and stability; 2) challenge-focused RL data derived from bad-case analysis further enhances perception and reasoning on difficult instances; and 3) training Qwen2.5-VL-7B with MHPR yields significant gains, achieving near-parity with considerably larger models. We release ACVG and MHPR to facilitate reproducible, extensible research on human-centric perception and reasoning.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.