HandFlow: 플로우 매칭을 이용한 완전 생성형 4차원 손 제스처 복원
HandFlow: Fully Generative 4D Hand Recovery with Flow Matching
정확한 단안(monocular) 4차원 손 제스처 복원은 여전히 어려운 과제입니다. 프레임별 판별 회귀 모델은 시간적 맥락이 부족하여 종종 불안정한 예측 결과를 초래합니다. 시간 모델은 여러 프레임을 통해 정보를 통합하여 일관성을 향상시키지만, 일반적으로 결정론적인 회귀 모델이기 때문에 가려짐(occlusion) 및 모션 블러로 인한 불확실한 관찰에 취약합니다. 생성 모델은 가능성 있는 손 제스처 시퀀스에 대한 사전 지식을 학습함으로써, 시각적 정보가 불완전하거나 신뢰할 수 없을 때 일관된 손 상태를 복원하는 자연스러운 대안을 제공합니다. 이러한 관찰에 따라, 본 논문에서는 시간적으로 일관성 있는 3차원 손 자세 및 형상 추정을 위한 완전 생성형 플로우 매칭 프레임워크인 HandFlow를 제안합니다. HandFlow는 시각적 및 골격 정보가 주어지면, 단일 ODE(Ordinary Differential Equation) 통합을 통해 전체 시간 창에 걸쳐 MANO 파라미터를 노이즈 제거합니다. 이를 지원하기 위해, 우리는 Flux 스타일의 이중 스트림 트랜스포머를 사용하여 자동 회귀 디코딩 없이 전체 시퀀스를 참조하여 장거리 의존성을 포착하고, 관찰된 특징을 학습 가능한 마스크 토큰과 혼합하는 신뢰도 기반 지속적 마스크 메커니즘을 활용하여 노이즈가 있거나 누락된 데이터를 처리합니다. DexYCB 및 HOT3D 데이터셋에 대한 실험 결과, HandFlow는 최첨단 성능을 달성했으며, 특히 공간 정확도 및 시간적 부드러움 측면에서 큰 개선을 보였습니다. HandFlow는 가장 강력한 기준 모델과 비교하여 공간 자세 오차를 30% 이상 줄이고, 평가된 모든 방법 중에서 가장 낮은 가속 오류를 달성하는 동시에 프레임별 자세 정확도에서도 경쟁력을 유지합니다. 또한, 단일 GPU 환경에서 HandFlow는 47fps의 속도로 150프레임 시퀀스를 복원하며, 이는 기존 비디오 기반 방법 중 가장 빠른 방법보다 약 12배 빠른 속도입니다. 이때 복원 자체는 전체 엔드-투-엔드 지연 시간의 작은 부분에 불과합니다.
Accurate monocular 4D hand reconstruction remains challenging. Per-frame discriminative regressors lack temporal context and often produce jittery predictions. Temporal models improve consistency by aggregating information across frames, but they are typically deterministic regressors, making them vulnerable to ambiguous observations caused by occlusion and motion blur. Generative modeling offers a natural alternative by learning a prior over plausible hand motion sequences, enabling coherent hand-state recovery when visual evidence is incomplete or unreliable. Motivated by this observation, we present HandFlow, a fully generative flow-matching framework for temporally coherent 3D hand pose and shape estimation from monocular video. Given visual and skeletal observations, HandFlow denoises an entire temporal window of MANO parameters through a single ODE integration. To support this, we use a Flux-style dual-stream transformer that attends across the full sequence to capture long-range dependencies without autoregressive decoding, and a confidence-aware continuous masking mechanism that blends observed features with learnable mask tokens to handle noisy or missing observations. Experiments on DexYCB and HOT3D show that HandFlow achieves state-of-the-art performance, with particularly large gains in world-space accuracy and temporal smoothness. It reduces world-space pose error by over 30% compared with the strongest baseline and achieves the lowest acceleration error among all evaluated methods, while remaining competitive in per-frame pose accuracy. Moreover, on a single GPU HandFlow reconstructs a 150-frame sequence at 47 fps, about 12x faster than the fastest prior video-based method, with reconstruction itself accounting for only a small fraction of the end-to-end latency.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.