강화 학습과 선택적 적대적 운동 사전 지식을 이용한 인간형 로봇의 다중 보행 학습
Multi-Gait Learning for Humanoid Robots Using Reinforcement Learning with Selective Adversarial Motion Prior
인간형 로봇의 다양한 보행 기술을 단일 강화 학습 프레임워크 내에서 학습하는 것은 다양한 보행 방식에 따른 안정성과 운동 성능 간의 상충되는 요구 사항으로 인해 여전히 어려운 과제입니다. 본 연구에서는 인간형 로봇이 일관된 정책 구조, 동작 공간 및 보상 체계를 사용하여 걷기, 오리 걸음, 달리기, 계단 오르기, 점프 등 다섯 가지 고유한 보행 방식을 익힐 수 있도록 하는 다중 보행 학습 접근 방식을 제시합니다. 핵심적인 기여는 선택적 적대적 운동 사전 지식(AMP) 전략입니다. AMP는 주기적이고 안정성이 중요한 보행 방식(걷기, 오리 걸음, 계단 오르기)에 적용되어 수렴 속도를 높이고 비정상적인 동작을 억제하며, 반면, 매우 역동적인 보행 방식(달리기, 점프)에서는 과도한 제약을 방지하기 위해 의도적으로 생략됩니다. 정책은 도메인 랜덤화를 적용한 시뮬레이션 환경에서 PPO를 사용하여 학습하고, 12 자유도 인간형 로봇에 제로샷 시뮬레이션-실제 환경 전이 방식을 통해 적용합니다. 정량적인 비교 결과, 선택적 AMP가 모든 다섯 가지 보행 방식에서 균일한 AMP 정책보다 우수한 성능을 나타냄을 보여줍니다. 선택적 AMP는 안정성에 중점을 둔 보행 방식에서 더 빠른 수렴 속도, 낮은 추적 오차 및 더 높은 성공률을 달성하는 동시에 역동적인 보행에 필요한 민첩성을 유지합니다.
Learning diverse locomotion skills for humanoid robots in a unified reinforcement learning framework remains challenging due to the conflicting requirements of stability and dynamic expressiveness across different gaits. We present a multi-gait learning approach that enables a humanoid robot to master five distinct gaits -- walking, goose-stepping, running, stair climbing, and jumping -- using a consistent policy structure, action space, and reward formulation. The key contribution is a selective Adversarial Motion Prior (AMP) strategy: AMP is applied to periodic, stability-critical gaits (walking, goose-stepping, stair climbing) where it accelerates convergence and suppresses erratic behavior, while being deliberately omitted for highly dynamic gaits (running, jumping) where its regularization would over-constrain the motion. Policies are trained via PPO with domain randomization in simulation and deployed on a physical 12-DOF humanoid robot through zero-shot sim-to-real transfer. Quantitative comparisons demonstrate that selective AMP outperforms a uniform AMP policy across all five gaits, achieving faster convergence, lower tracking error, and higher success rates on stability-focused gaits without sacrificing the agility required for dynamic ones.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.