2606.09559v1 Jun 08, 2026 cs.LG

안전 강화 학습의 지식 제거: Safe-RULE

Safe-RULE: Safe Reinforcement UnLEarning

Fanxin Kong
Fanxin Kong
Citations: 75
h-index: 5
Shixiong Jiang
Shixiong Jiang
Citations: 12
h-index: 2
Taozheng Zhu
Taozheng Zhu
Citations: 10
h-index: 2

오프라인 안전 강화 학습(Safe RL)은 온라인 상호작용 없이 정책을 학습할 수 있도록 하여, 로봇 시스템과 같은 안전이 중요한 시스템에 적합합니다. 그러나 오프라인 Safe RL은 정적인 데이터셋에 의존하기 때문에 데이터 포이즈닝 공격에 취약하며, 이러한 공격에서 악의적인 샘플이 주입되어 안전성을 저해하고 위험한 정책을 유발할 수 있습니다. 본 연구에서는 안전 강화 학습의 지식 제거(Safe-RULE)라는 새로운 학습 패러다임을 제안합니다. 이는 방어 프레임워크로 사용되며, 원본 학습 환경에 대한 접근 없이 또는 처음부터 다시 훈련하지 않고도 악성 데이터의 영향을 제거합니다. 또한, 우리는 작업 성능과 안전 제약 조건을 모두 고려하여 오프라인 Safe RL에 강화 학습 지식 제거를 확장했습니다. 벤치마크 Safe RL 작업에서의 실험 결과는 우리 방법이 데이터 포이즈닝 공격에 대한 안전 성능을 효과적으로 향상시킨다는 것을 보여줍니다.

Original Abstract

Offline safe reinforcement learning (Safe RL) enables policy learning without online interactions, making it suitable for safety-critical systems such as robotics systems. However, its reliance on static datasets exposes offline Safe RL to data poisoning attacks, where adversaries inject malicious samples that compromise safety and induce unsafe policy behavior. In this work, we propose a new learning paradigm, named safe reinforcement unlearning (Safe-RULE), used as a defense framework to remove the influence of poisoned data without retraining from scratch or requiring access to the original training environment. We further extend reinforcement unlearning to offline Safe RL by explicitly accounting for both task performance and safety constraints during the unlearning process. Experiments across benchmark Safe RL tasks demonstrate that our approach effectively enhances safety performance against data poisoning attacks.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!