2606.25680v1 Jun 24, 2026 cs.RO

제약 조건 강화 학습을 이용한 전력 예산 기반 수중 로봇 제어

Power-Budgeted Underwater Vehicle Control via Constrained Reinforcement Learning

Yinuo Wang
Yinuo Wang
Citations: 229
h-index: 7
Gavin Tao
Gavin Tao
Citations: 7
h-index: 1
John V. Ringwood
John V. Ringwood
Citations: 0
h-index: 0
Yuze Liu
Yuze Liu
Citations: 0
h-index: 0

수중 로봇은 제한된 온보드 에너지 예산을 사용하며, 추진력으로 인해 에너지가 빠르게 소모됩니다. 따라서, 주어진 작업을 수행하면서 더 적은 추력 전력을 사용하는 컨트롤러는 임무 범위와 지속 시간을 직접적으로 늘립니다. 강화 학습은 정박 및 경로 추적을 위한 효과적인 모델 기반 제어기를 제공하지만, 작업 정확도만 최적화하면 정책이 진동하는 방식으로 작동하여 에너지를 낭비하게 됩니다. 기존의 해결책은 보상에서 에너지 페널티를 빼는 것이지만, 이는 물리적 단위가 없는 단일 가중치를 사용하여 작업-전력 간의 균형을 설정합니다. 즉, 목표 전력 수준을 지정할 수 없으며, 가중치는 각 로봇과 작업에 대해 재조정해야 하며, 잘못된 가중치는 오히려 전력을 증가시킬 수도 있습니다. 본 논문에서는 에너지 효율적인 수중 제어를 평균 추력 전력이 명시적인 예산 내에 있도록 제한되는 제약 조건 마르코프 결정 프로세스로 공식화하고, PPO-라그랑주 알고리즘을 사용하여 이를 해결합니다. 전력 수준은 물리적 단위로 예산을 설정하여 조정하며, 단일 이중 변수를 온라인으로 업데이트하여 각 로봇과 작업에 대해 해당 예산 조건을 충족하도록 합니다. MarineGym 시뮬레이터에서 세 대의 로봇과 네 가지 작업을 수행한 결과, 에너지 제약 정책은 모든 12가지 설정에서 가장 적은 전력을 소비했으며, 작업 전용 기준보다 14~65% (최대 64.9%) 절감 효과를 보였으며, 에너지-보상 기준보다 항상 우수했습니다. 또한, 대부분의 경우 부드러운 작동을 유지하며, 하나의 의도적으로 전력 제한이 설정된 환경을 제외하고는 작업 정확도를 유지했습니다. 따라서, 에너지를 명시적인 제약 조건으로 부과하면 로봇 및 작업별 가중치 검색 없이 에너지 효율적인 수중 제어를 구현할 수 있는 방법을 제공합니다.

Original Abstract

Underwater vehicles operate from a fixed onboard energy budget that propulsion rapidly depletes, so a controller that completes its task while drawing less thruster power directly extends mission range and endurance. Reinforcement learning yields capable model-free controllers for station-keeping and trajectory tracking, but optimizing task accuracy alone drives the policy toward oscillatory, energy-wasting actuation. The established remedy subtracts an energy penalty from the reward, yet this sets the task-power trade-off through a single weight with no physical units: a target power level cannot be specified, the weight must be re-tuned for every vehicle and task, and a mismatched weight can even raise power. This paper instead formulates energy-efficient underwater control as a constrained Markov decision process in which average thruster power is subject to an explicit budget, solved with a PPO-Lagrangian algorithm. The power level is set by declaring a budget in physical units, and a single dual variable is updated online to meet it for each vehicle and task, without manual weight search. Across three vehicles and four tasks in the MarineGym simulator, the energy-constrained policy draws the least power in all twelve settings, reducing it by 14--65\% (up to 64.9\%) over a task-only baseline and below an energy-reward baseline everywhere, while remaining the smoothest in ten settings and preserving task accuracy except in one deliberately power-limited regime. Imposing energy as an explicit constraint thus offers a tuning-free route to energy-efficient underwater control that needs no per-vehicle, per-task weight search.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!