숙련된 조작을 위한 최소한의 리타겟팅 기반 강화 학습 방법론
A Minimalist Retargeting-Guided Reinforcement Learning Recipe for Dexterous Manipulation
최근 휴머노이드 로봇 전체 제어 분야에서, 인간 동작을 로봇 운동학 참조점으로 변환하고, 강화 학습(RL)을 통해 이를 추종하도록 정책을 훈련하는 간단한 방법이 성공적인 결과를 보여주었습니다. 하지만 이 방법은 숙련된 조작에 어떻게 적용될까요? 조작은 복잡하고 접촉이 많은 동역학을 포함하며, 미세한 접촉 모드 및 힘 제어가 필요하기 때문에 명확하지 않습니다. 본 논문에서는 단일 인간 시연으로부터 숙련된 조작 정책을 학습하는 최소한의 리타겟팅 기반 RL 파이프라인인 REGRIND를 제시합니다. REGRIND는 인간의 손-물체 운동을 로봇 참조점으로 변환하여, 손과 물체의 공간적 관계 및 접촉 관계를 유지하고, 시뮬레이션 환경에서 해당 참조점을 따라 물체를 중심으로 하는 주요 지점을 추종하는 잔차 RL 정책을 훈련하며, 신중한 시스템 식별 과정을 통해 학습된 정책을 실제 하드웨어로 즉시 이전합니다. 결과적으로 얻어진 정책은 두 가지 다른 다관절 손에 대해 접촉이 많은 도구 사용 작업에서 부드럽고 인간과 유사한 동작을 생성합니다. 체계적인 하드웨어 실험을 통해, 숙련된 조작에서의 시뮬레이션-실제 이전(sim-to-real transfer)을 지배하는 주요 요인을 식별하고 분석하며, 접촉이 많은 환경에서의 리타겟팅 기반 학습에 대한 실질적인 가이드라인을 제공합니다. 동영상 및 코드는 https://yunhaifeng.com/REGRIND 에서 확인할 수 있습니다.
Recent work in humanoid whole-body control has found success with a simple recipe: retarget human motion to robot kinematic references, then train policies via reinforcement learning (RL) to track them. But how does this recipe transfer to dexterous manipulation? The answer is not obvious, as manipulation involves complex, contact-rich dynamics and requires delicate regulation of contact modes and forces. We present REGRIND, a minimalist retargeting-guided RL pipeline that learns dexterous manipulation policies from a single human demonstration. REGRIND retargets human hand-object motion to a robot reference that preserves hand-object spatial and contact relationships, trains a residual RL policy in simulation to track object-centric keypoints along that reference, and transfers the resulting policy zero-shot to hardware with careful system identification. The resulting policies produce fluid, human-like behavior on two different multi-fingered hands across contact-rich tool-use tasks, including operating a pair of scissors and turning a screwdriver. Through systematic hardware experiments, we identify and analyze the key factors that govern sim-to-real transfer in dexterous manipulation, offering practical guidance for retargeting-based learning in contact-rich settings. Videos and code are available at https://yunhaifeng.com/REGRIND.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.