CoorDex: 연속적인 숙련형 인간형 로봇의 보행 및 조작을 위한 신체 및 손 제어 우선순위 조정
CoorDex: Coordinating Body and Hand Priors for Continuous Dexterous Humanoid Loco-Manipulation
인간형 로봇의 보행 및 조작은 종종 멈춤-조작-재개 과정으로 단순화됩니다. 또한, 이는 일반적으로 개방-닫힘 동작과 유사한 낮은 자유도(DoF) 엔드 이펙터를 사용합니다. 본 연구에서는 CoorDex라는 학습 파이프라인을 소개합니다. CoorDex는 고차원적인 신체 및 숙련형 손 제어를 조정된 잠재 잔차 제어로 변환하여, 로봇이 움직이는 동안에도 높은 자유도를 가진 숙련형 조작을 가능하게 합니다. CoorDex는 시뮬레이션 환경에서 얻은 전체 신체 및 손 동작 데이터를 기반으로, 인간형 로봇의 신체와 숙련형 손에 대한 모션 추적 모델을 학습하고, 이를 고유 수용성 감각에 조건화된 잠재 우선순위로 변환합니다. 이러한 우선순위를 강화학습의 행동 공간으로 사용하여 잔차 강화를 수행합니다. 조정된 잠재 잔차 정책은 공유된 작업 컨텍스트와 별도의 신체-손 잔차 네트워크를 통해 이러한 우선순위를 결합하여 자연스러운 전체 신체 동작을 유지하면서 손가락 수준의 접촉 안정성을 향상시킵니다. CoorDex는 20 자유도를 가진 WUJI 핸드를 장착한 Unitree G1 인간형 로봇이 움직이는 동안 숙련형 조작을 수행할 수 있도록 합니다. 예를 들어, 멈추지 않고 병을 잡고 운반하거나, 이동하면서 냉장고 문을 여는 동작 등이 가능합니다.
Humanoid loco-manipulation is often simplified into a stop-and-go process: walking to an object, stopping to manipulate it, and then resuming locomotion. It also commonly relies on low degree-of-freedom (DoF) end effectors that behave like an open-close grasp primitive. We introduce CoorDex, a learning pipeline that converts high-dimensional body and dexterous hand control into coordinated latent residual control, enabling high-DoF dexterous loco-manipulation on the move. Starting from simulated whole-body and hand demonstrations, CoorDex trains privileged motion tracking teachers for the humanoid body and dexterous hand, distills them into proprioception-conditioned latent priors, and uses the frozen priors as the action space for downstream residual reinforcement learning. A coordinated latent residual policy composes these priors through shared task context and separate body-hand residual heads, preserving natural whole-body motion while improving finger-level contact reliability. CoorDex enables a Unitree G1 humanoid with a 20-DoF WUJI hand to execute dexterous manipulation while in motion, including non-stop bottle grasping and carrying, fridge door opening on the move, and cube pick-and-turn. Ablations on the walk-grasp-carry task show that joint-space PPO, joint-space hand control, and monolithic latent prediction all fail under the same reward budget, while the latent-prior interface and coordinated residual structure make high-dimensional contact-rich loco-manipulation trainable. Project Page: https://skevinci.github.io/coordex/
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.