2607.00442v1 Jul 01, 2026 cs.RO

시계열 논리 명세를 이용한 보행 인식 사족 보행 제어 학습

Learning Gait-Aware Quadruped Locomotion with Temporal Logic Specifications

Merve Atasever
Merve Atasever
Citations: 0
h-index: 0
Cagan Bakirci
Cagan Bakirci
Citations: 0
h-index: 0
Alfredo Reina Corona
Alfredo Reina Corona
Citations: 0
h-index: 0
Keyan Azbijari
Keyan Azbijari
Citations: 1
h-index: 1
Jyotirmoy V. Deshmukh
Jyotirmoy V. Deshmukh
Citations: 3,596
h-index: 30

사족 로봇의 보행 제어를 위한 강화학습(RL)은 일반적으로 고정되고, 사람이 직접 설계하며, 마르코프 성질을 갖는 보상 함수에 의존하는데, 이는 학습된 정책의 해석 가능성을 제한하고 보행 동작에 대한 명시적인 제어를 제공하지 못합니다. 본 연구에서는 Signal Temporal Logic (STL)으로 표현된 매개변수화된 제약 조건을 사용하여 다양한 보행 패턴을 정의하는 프레임워크를 제시합니다. 이러한 제약 조건에는 안전 범위, 보행 동기화 제약, 명령 추적 및 작동 제한이 포함됩니다. 이러한 명세를 기반으로, 학습 에이전트에게 원하는 동작을 인코딩하는 밀집되고 연속적인 보상 지형을 제공하는 보상 설계 메커니즘을 개발했습니다. 세 가지 속도 범위(보행-조깅, 조깅, 전력 질주)에 대한 매개변수화된 STL 템플릿을 정의하고, 기준 실행 결과를 사용하여 해당 매개변수를 조정하며, STL의 강건성을 부드럽게 근사하여 보상을 계산합니다. 생성된 보상은 Proximal Policy Optimization (PPO)와 호환되는 형상화된 기울기를 제공하는 데 사용될 수 있습니다. 본 연구에서는 Google의 Barkour 사족 로봇을 MuJoCo XLA (MJX) 환경에서 사용하여 제안하는 방법을 구현했습니다. 시뮬레이터 내의 병렬화를 통해 학습 속도를 향상시키고, 도메인 랜덤화를 사용하여 학습된 정책의 강건성을 확보했습니다. 실험 결과, 사람이 직접 설계한 보상에 비해 STL 기반으로 형상화된 보상이 더 정확한 속도 추적 성능과 안정적인 학습 결과를 제공하는 것을 확인했습니다. 관련 동영상은 프로젝트 웹사이트에서 확인할 수 있습니다: https://stl-locomotion.github.io/.

Original Abstract

Reinforcement learning (RL) for quadruped locomotion commonly depends on fixed, hand-crafted, and Markovian reward functions that limit both interpretability of learned policies and lack explicit control over gait behaviors. We introduce a framework where distinct gaits are specified using parameterized constraints expressed in Signal Temporal Logic (STL). These include safety bounds, gait synchronization constraints, command tracking, and actuation bounds. From these specifications, we develop a reward shaping mechanism that provides learning agents a dense, continuous reward landscape that encodes desired behavior. We define parametric STL templates for three speed regimes (walking-trot, trot, bound), calibrate their parameters from reference rollouts, and compute rewards from using smooth approximations of STL robustness over the rollouts. The generated rewards can be used to provide shaped gradients compatible with Proximal Policy Optimization (PPO). We instantiate the approach on Google's Barkour quadruped robot in MuJoCo XLA (MJX). We use parallelization within the simulator to improve training speeds and use domain randomization to robustify learned policies. We show that compared to a baseline of hand-crafted rewards, the STL-shaped rewards yield tighter velocity tracking and more stable training. Videos can be found on our project website: https://stl-locomotion.github.io/.

0 Citations
0 Influential
15 Altmetric
75.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!