2608.03158v1 Aug 04, 2026 cs.CV

다중 객체 및 관절형 인간-객체 상호 작용 생성을 위한 표면 키포인트 표현

Surface Keypoint Representation for Multi-Object and Articulated Human-Object Interaction Generation

Jihua Zhu
Jihua Zhu
Citations: 3,102
h-index: 28
Xiaogang Peng
Xiaogang Peng
Citations: 87
h-index: 2
Zeyu Han
Zeyu Han
Citations: 1,236
h-index: 8
Zichong Meng
Zichong Meng
Citations: 87
h-index: 2
Yiming Xie
Yiming Xie
Citations: 1,341
h-index: 13
Gang Hua
Gang Hua
Citations: 419
h-index: 11
Huaizu Jiang
Huaizu Jiang
Citations: 745
h-index: 11

일상적인 활동은 인간이 전신 동작을 주변 객체의 동작과 조화시키는 것을 필요로 합니다. 최근 인간-객체 상호 작용 (HOI) 생성 분야에서 상당한 진전이 있었지만, 대부분의 기존 방법은 단일의 강체 객체와의 상호 작용을 가정하며, 가변적인 수의 객체 또는 다양한 관절 메커니즘을 가진 객체가 관련된 시나리오에는 잘 적용되지 않습니다. 본 연구에서는 객체 동작 표현으로 표면 키포인트 궤적을 제안합니다. 각 강체 구성 요소(독립적인 객체이든 복잡한 조립체의 일부이든)에 대해, 시간에 따른 작은 수의 비평행 표면 점들을 추적합니다. 이러한 표현 방식은 명시적인 관절 유형 사양 없이 점 역학을 통해 다중 객체 조정 및 다양한 관절 메커니즘을 직접적으로 처리할 수 있습니다. 각 신체 부위가 어떤 객체와 언제, 어디에서 접촉하는지를 모델링하기 위해, 우리는 거리 기반 접촉 모델링을 전신, 다중 객체, 그리고 관절형 환경으로 확장하는 시공간 접촉 거리장을 도입합니다. HOI 생성을 세 단계로 나누어 진행합니다: 텍스트 또는 웨이포인트로부터 객체 동작을 생성하고, 접촉 거리장을 예측하며, 접촉-유도 최적화를 통해 전신 동작을 합성합니다. ParaHome, HIMO, ARCTIC 및 OMOMO 데이터셋에서의 실험 결과는 단일 객체, 다중 객체 및 관절형 상호 작용 환경에서 기존 방법보다 더 나은 또는 동등한 성능을 보여줍니다.

Original Abstract

Daily activities require humans to coordinate whole-body motion with the motion of surrounding objects. Despite recent progress in human-object interaction (HOI) generation, most existing methods assume interactions with a single rigid object and do not extend well to scenarios involving a variable number of objects or articulated objects with diverse joint mechanisms. We propose surface keypoint trajectories as an object motion representation: for each rigid component, whether a standalone object or one part of an articulated assembly, we track a small set of non-collinear surface points over time. This representation handles multi-object coordination and diverse articulation mechanisms directly from point dynamics without requiring explicit joint-type specification. To model when and where each body region contacts each object, we introduce a spatio-temporal contact distance field that extends distance-based contact modeling to whole-body, multi-object, and articulated settings. We factorize HOI generation into three stages: generating object motions from text or waypoints, predicting the contact distance field, and synthesizing whole-body motion with contact-guided optimization. Experiments on ParaHome, HIMO, ARCTIC, and OMOMO demonstrate better or comparable performance to existing methods across single-object, multi-object, and articulated interaction settings.

0 Citations
0 Influential
14 Altmetric
70.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!