2604.23576v1 Apr 26, 2026 cs.LG

CAPSULE: 제어 이론 기반 액션 섭동을 활용한 안전하고 불확실성-인지 강화 학습

CAPSULE: Control-Theoretic Action Perturbations for Safe Uncertainty-Aware Reinforcement Learning

O. Jain
O. Jain
Citations: 3
h-index: 1
Rahul Narava
Rahul Narava
Citations: 1
h-index: 1
Siddharth Verma
Siddharth Verma
Citations: 127
h-index: 5
S. Jha
S. Jha
Citations: 269
h-index: 9
M. Jha
M. Jha
Citations: 507
h-index: 12

알 수 없는 동역학을 가진 고차원 시스템에서 안전한 탐색을 보장하는 것은 여전히 중요한 과제입니다. 기존의 안전 강화 학습 방법은 종종 기대값 수준에서만 안전성을 보장하며, 이로 인해 안전 위반이 발생할 수 있습니다. 반면, 제어 이론 기반 접근 방식은 엄격한 제약 조건 기반의 안전성을 제공하지만, 일반적으로 알려진 시스템 동역학에 대한 접근성을 필요로 하거나, 제어 친화형 모델의 정확한 추정을 요구합니다. 본 논문에서는 오프라인 환경에서 확률적 제어 친화형 동역학 모델을 학습하는 안전 강화 학습 프레임워크를 제안합니다. 학습된 모델을 활용하여 모델 불확실성을 고려하여 보수적인 안전 제약 조건을 제공하는 제어 장벽 함수(CBF)를 명시적으로 구성합니다. 이러한 CBF 제약 조건은 온라인 기반의 제약 조건 부재 액션 수정 메커니즘을 통해 적용되어, 과도하게 작업 성능을 제한하지 않고 안전한 탐색을 가능하게 합니다. 비선형, 복잡한 연속 제어 벤치마크에 대한 실험적 평가 결과, 제안하는 방법은 기존 방법과 유사한 보상을 달성하면서 안전 위반을 크게 줄이는 것을 보여줍니다.

Original Abstract

Ensuring safe exploration in high-dimensional systems with unknown dynamics remains a significant challenge. Existing safe reinforcement learning methods often provide safety guarantees only in expectation, which can still lead to safety violations. Control-theoretic approaches, in contrast, offer hard constraint-based safety guarantees but typically assume access to known system dynamics or require accurate estimation of control-affine models. In this paper, we propose a safe reinforcement learning framework that learns a probabilistic control-affine dynamics model in an offline setting. The learned model is leveraged to explicitly construct control barrier functions (CBFs) that incorporate model uncertainty to provide conservative safety constraints. These CBF constraints are enforced through an online constraint-based action correction mechanism, enabling safe exploration without overly restricting task performance. Empirical evaluations on nonlinear, complex continuous-control benchmarks demonstrate that our approach achieves returns comparable to those of existing baselines while significantly reducing safety violations.

0 Citations
0 Influential
6 Altmetric
30.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!