다중 객체 및 관절형 인간-객체 상호 작용 생성을 위한 표면 키포인트 표현
Surface Keypoint Representation for Multi-Object and Articulated Human-Object Interaction Generation
일상적인 활동은 인간이 전신 동작을 주변 객체의 동작과 조화시키는 것을 필요로 합니다. 최근 인간-객체 상호 작용 (HOI) 생성 분야에서 상당한 진전이 있었지만, 대부분의 기존 방법은 단일의 강체 객체와의 상호 작용을 가정하며, 가변적인 수의 객체 또는 다양한 관절 메커니즘을 가진 객체가 관련된 시나리오에는 잘 적용되지 않습니다. 본 연구에서는 객체 동작 표현으로 표면 키포인트 궤적을 제안합니다. 각 강체 구성 요소(독립적인 객체이든 복잡한 조립체의 일부이든)에 대해, 시간에 따른 작은 수의 비평행 표면 점들을 추적합니다. 이러한 표현 방식은 명시적인 관절 유형 사양 없이 점 역학을 통해 다중 객체 조정 및 다양한 관절 메커니즘을 직접적으로 처리할 수 있습니다. 각 신체 부위가 어떤 객체와 언제, 어디에서 접촉하는지를 모델링하기 위해, 우리는 거리 기반 접촉 모델링을 전신, 다중 객체, 그리고 관절형 환경으로 확장하는 시공간 접촉 거리장을 도입합니다. HOI 생성을 세 단계로 나누어 진행합니다: 텍스트 또는 웨이포인트로부터 객체 동작을 생성하고, 접촉 거리장을 예측하며, 접촉-유도 최적화를 통해 전신 동작을 합성합니다. ParaHome, HIMO, ARCTIC 및 OMOMO 데이터셋에서의 실험 결과는 단일 객체, 다중 객체 및 관절형 상호 작용 환경에서 기존 방법보다 더 나은 또는 동등한 성능을 보여줍니다.
Daily activities require humans to coordinate whole-body motion with the motion of surrounding objects. Despite recent progress in human-object interaction (HOI) generation, most existing methods assume interactions with a single rigid object and do not extend well to scenarios involving a variable number of objects or articulated objects with diverse joint mechanisms. We propose surface keypoint trajectories as an object motion representation: for each rigid component, whether a standalone object or one part of an articulated assembly, we track a small set of non-collinear surface points over time. This representation handles multi-object coordination and diverse articulation mechanisms directly from point dynamics without requiring explicit joint-type specification. To model when and where each body region contacts each object, we introduce a spatio-temporal contact distance field that extends distance-based contact modeling to whole-body, multi-object, and articulated settings. We factorize HOI generation into three stages: generating object motions from text or waypoints, predicting the contact distance field, and synthesizing whole-body motion with contact-guided optimization. Experiments on ParaHome, HIMO, ARCTIC, and OMOMO demonstrate better or comparable performance to existing methods across single-object, multi-object, and articulated interaction settings.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.