2607.13646v1 Jul 15, 2026 cs.CV

Human4K: 전신 3차원 인간 재구성을 위한 고해상도 다중 시점 모션 캡처 데이터셋

Human4K: A Large-Scale 4K Multi-View Mocap Dataset for Whole-Body 3D Human Reconstruction

Tianshun Han
Tianshun Han
Citations: 23
h-index: 3
Ziyu Shi
Ziyu Shi
Citations: 0
h-index: 0
Lijiang Liu
Lijiang Liu
Citations: 13
h-index: 2
Ajian Liu
Ajian Liu
Citations: 1,724
h-index: 20
Benjia Zhou
Benjia Zhou
Citations: 531
h-index: 10
H. Escalante
H. Escalante
Citations: 192
h-index: 3
Yanyan Liang
Yanyan Liang
Citations: 1,052
h-index: 14
Sergio Escalera
Sergio Escalera
Citations: 97
h-index: 4
Zhen Lei
Zhen Lei
Citations: 499
h-index: 11
Jun Wan
Jun Wan
Citations: 4,139
h-index: 32

최근 3차원 인간 재구성 기술의 발전으로 전체적인 성능이 향상되었지만, 현재 모델은 여전히 가장 어려운 실제 환경에서 어려움을 겪고 있습니다. 이러한 모델들은 종종 불안정한 기하 구조를 생성하고, 부정확한 사지 관절 움직임을 나타내며, 깊이 불명확성 또는 자기 가림 현상이 발생하는 경우 신뢰할 수 없는 예측을 합니다. 이는 기존 데이터셋이 강력한 재구성을 지원하기 위해 필요한 고해상도 이미지, 고정밀 어노테이션 및 다양한 전신 동작의 조합이 부족하기 때문입니다. 이러한 문제를 해결하기 위해, 우리는 SMPL-X 어노테이션으로 정확하게 모션 캡처된, 대규모 4K 다중 시점 전신 인간 재구성 데이터셋인 Human4K를 소개합니다. Human4K는 전문적인 Vicon 모션 캡처 시스템과 동기화된 8개의 고해상도 카메라 시스템을 통해 촬영된 6백만 장 이상의 4K 이미지를 포함하며, 11명의 피험자가 복잡하고 정교한 전신 동작을 수행하는 모습을 담고 있습니다. 모든 시퀀스는 전체 몸통 및 사지의 정확한 정렬을 보장하기 위해 모션 재매핑 및 개선 모듈(MRRM)을 통해 처리되었습니다. 실험 결과는 Human4K를 사용하여 학습하면 표준 벤치마크에서 전신 재구성이 꾸준히 향상되며, 특히 손, 발 및 깊이 불명확한 사지 구성에 대해 상당한 성능 향상을 보인다는 것을 보여줍니다.

Original Abstract

Recent advances in 3D human reconstruction have improved overall performance, yet current models still fail in the most challenging real-world scenarios. They often produce unstable geometry, inaccurate limb articulation and unreliable predictions under depth ambiguity or self-occlusion. A key reason is that existing datasets still lack the combination of high-resolution images, high-precision annotations and diverse whole-body motions required to support robust reconstruction. To address this gap, we present Human4K, a large-scale 4K multi-view whole-body human reconstruction dataset with mocap-accurate SMPL-X annotations. Human4K contains over six million 4K images captured by an eight-view high-resolution camera system synchronized with a professional Vicon motion capture setup, covering 11 subjects performing complex, highly articulated and strongly self-occluded full-body motions. All sequences are processed by a Motion-Retargeting and Refinement Module (MRRM) to ensure precise alignment for the full body and extremities. Experimental results show that training with Human4K consistently improves whole-body reconstruction on standard benchmarks, with particularly large gains for hands, feet and depth-ambiguous limb configurations.

0 Citations
0 Influential
16 Altmetric
80.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!