2604.19104v1 Apr 21, 2026 cs.RO

강화 학습 기반 적응형 다중 작업 제어를 통한 이족 로봇 축구 로봇 제어

Reinforcement Learning Enabled Adaptive Multi-Task Control for Bipedal Soccer Robots

Linqi Ye
Linqi Ye
Citations: 48
h-index: 3
Ting Wu
Ting Wu
Citations: 31
h-index: 3
Yulai Zhang
Yulai Zhang
Citations: 19
h-index: 2
Yinrong Zhang
Yinrong Zhang
Citations: 0
h-index: 0

동적 환경에서 작동하는 이족 축구 로봇 개발은 운동 안정성과 여러 작업 간의 복잡한 연관성, 그리고 보행 및 낙상 복구와 같은 다양한 상태 간의 제어 전환 문제 등 여러 가지 어려움을 야기합니다. 이러한 문제점을 해결하기 위해, 본 논문에서는 적응형 다중 작업 제어를 달성하기 위한 모듈형 강화 학습(RL) 프레임워크를 제안합니다. 먼저, 본 프레임워크는 개방 루프 피드포워드 발진기와 강화 학습 기반 피드백 잔차 전략을 결합하여, 기본적인 보행 패턴 생성과 복잡한 축구 동작을 효과적으로 분리합니다. 둘째, 자세 기반 상태 머신을 도입하여, 공 탐색 및 킥 네트워크(BSKN)와 낙상 복구 네트워크(FRN) 간의 명확한 전환을 통해 상태 간섭을 근본적으로 방지합니다. FRN은 점진적인 힘 감쇠 커리큘럼 학습 전략을 통해 효율적으로 훈련됩니다. 제안된 아키텍처는 Unity 시뮬레이션을 통해 이족 로봇에 적용되었으며, 뛰어난 공간 적응성(제한된 코너 상황에서도 안정적으로 공을 찾고 킥하는 능력)과 빠른 자율적 낙상 복구(평균 복구 시간 0.715초)를 보여주었습니다. 이는 복잡한 다중 작업 환경에서 원활하고 안정적인 작동을 보장합니다.

Original Abstract

Developing bipedal football robots in dynamiccombat environments presents challenges related to motionstability and deep coupling of multiple tasks, as well ascontrol switching issues between different states such as up-right walking and fall recovery. To address these problems,this paper proposes a modular reinforcement learning (RL)framework for achieving adaptive multi-task control. Firstly,this framework combines an open-loop feedforward oscilla-tor with a reinforcement learning-based feedback residualstrategy, effectively separating the generation of basic gaitsfrom complex football actions. Secondly, a posture-driven statemachine is introduced, clearly switching between the ballseeking and kicking network (BSKN) and the fall recoverynetwork (FRN), fundamentally preventing state interference.The FRN is efficiently trained through a progressive forceattenuation curriculum learning strategy. The architecture wasverified in Unity simulations of bipedal robots, demonstratingexcellent spatial adaptability-reliably finding and kicking theball even in restricted corner scenarios-and rapid autonomousfall recovery (with an average recovery time of 0.715 seconds).This ensures seamless and stable operation in complex multi-task environments.

0 Citations
0 Influential
1.5 Altmetric
7.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!