2604.03299v1 Mar 29, 2026 cs.CV

MoViD: 모션-뷰 분리 기반 시점 불변 3차원 인간 자세 추정

MoViD: View-Invariant 3D Human Pose Estimation via Motion-View Disentanglement

Hengle Jiang
Hengle Jiang
Citations: 227
h-index: 5
Yejia Liu
Yejia Liu
Citations: 2
h-index: 1
Haoxiang Liu
Haoxiang Liu
Citations: 4
h-index: 1
Runxi Huang
Runxi Huang
Citations: 4
h-index: 1
Xiaomin Ouyang
Xiaomin Ouyang
Citations: 803
h-index: 11

3차원 인간 자세 추정은 의료 모니터링, 인간-로봇 협업, 몰입형 게임 등 다양한 분야에서 핵심적인 기술이지만, 실제 적용에서는 시점 변화로 인한 어려움이 존재합니다. 기존 방법들은 새로운 카메라 시점에 대한 일반화 성능이 낮고, 많은 양의 학습 데이터가 필요하며, 높은 추론 지연 시간을 갖습니다. 본 논문에서는 시점 정보를 모션 특징과 분리하여 시점 불변적인 3차원 인간 자세 추정 프레임워크인 MoViD를 제안합니다. 핵심 아이디어는 중간 자세 특징에서 시점 정보를 추출하고, 이를 활용하여 자세 추정의 견고성과 효율성을 향상시키는 것입니다. MoViD는 주요 관절 관계를 모델링하여 시점 정보를 예측하는 시점 추정 모듈과, 모션 및 시점 특징을 분리하는 직교 투영 모듈을 도입합니다. 또한, 물리 법칙 기반의 대비 학습을 통해 시점 간의 특징을 정렬하여 성능을 더욱 향상시켰습니다. 실시간 엣지 환경에서의 적용을 위해, MoViD는 시점 정보를 고려한 전략을 통해 프레임 단위로 추론을 수행하며, 추정된 시점에 따라 플립 정제 과정을 적응적으로 활성화합니다. 9개의 공개 데이터셋과 새로 수집된 다중 시점 드론 및 보행 분석 데이터셋에 대한 평가 결과, MoViD는 최첨단 방법보다 자세 추정 오류를 24.2% 이상 감소시키고, 60% 적은 학습 데이터로도 견고한 성능을 유지하며, NVIDIA 엣지 장치에서 15 FPS의 실시간 추론 속도를 달성합니다.

Original Abstract

3D human pose estimation is a key enabling technology for applications such as healthcare monitoring, human-robot collaboration, and immersive gaming, but real-world deployment remains challenged by viewpoint variations. Existing methods struggle to generalize to unseen camera viewpoints, require large amounts of training data, and suffer from high inference latency. We propose MoViD, a viewpoint-invariant 3D human pose estimation framework that disentangles viewpoint information from motion features. The key idea is to extract viewpoint information from intermediate pose features and leverage it to enhance both the robustness and efficiency of pose estimation. MoViD introduces a view estimator that models key joint relationships to predict viewpoint information, and an orthogonal projection module to disentangle motion and view features, further enhanced through physics-grounded contrastive alignment across views. For real-time edge deployment, MoViD employs a frame-by-frame inference pipeline with a view-aware strategy that adaptively activates flip refinement based on the estimated viewpoint. Evaluations on nine public datasets and newly collected multiview UAV and gait analysis datasets show that MoViD reduces pose estimation error by over 24.2\% compared to state-of-the-art methods, maintains robust performance under severe occlusions with 60\% less training data, and achieves real-time inference at 15 FPS on NVIDIA edge devices.

1 Citations
0 Influential
5.5 Altmetric
28.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!