2607.21111v1 Jul 23, 2026 cs.LG

TOUR: 오프라인 강화 학습을 위한 경로 수준의 데이터 삭제 성능 평가 기준

TOUR: A Trajectory-Level Unlearning Benchmark for Offline Reinforcement Learning

Xin Yang
Xin Yang
Citations: 27
h-index: 2
Lingfei Ren
Lingfei Ren
Citations: 10
h-index: 2
Chaofan Pan
Chaofan Pan
Citations: 37
h-index: 4
Xiangyu Jiang
Xiangyu Jiang
Citations: 49
h-index: 2
Xuemei Cao
Xuemei Cao
Citations: 83
h-index: 4
Wei Wei
Wei Wei
Citations: 23
h-index: 2
Yanhua Li
Yanhua Li
Citations: 101
h-index: 5
Hao Yu
Hao Yu
Citations: 260
h-index: 6
Xiangkun Wang
Xiangkun Wang
Citations: 67
h-index: 4

오프라인 강화 학습(RL) 에이전트는 고정된 행동 경로 데이터를 기반으로 학습되므로, 학습 후 특정 데이터를 제거해야 할 경우 경로 단위로 데이터를 삭제하는 것이 중요합니다. 그러나 이러한 삭제 과정을 평가하기는 어렵습니다. 낮은 멤버십 점수는 경로 삭제를 반영할 수도 있지만, 다른 공격에 의해 남은 데이터가 기억될 수도 있고, 유용한 행동을 파괴하는 정책 붕괴를 초래할 수도 있기 때문입니다. 본 논문에서는 오프라인 RL에서 경로 수준의 기억 및 데이터 삭제 성능을 평가하기 위한 기준인 Trajectory-level memOrization and Unlearning (TOUR)을 제안합니다. TOUR은 경로 단위 분할, 매칭된 비멤버 컨트롤 그룹, 재학습 참조 데이터, 유지된 성능 지표, 그리고 다중 공격 기반 개인 정보 감사 기능을 결합합니다. D4RL 로봇 움직임 실험과 탐색형 AntMaze 확장에 대한 실험 결과, 일반적인 삭제 방법들이 환경에 따라 개인 정보 보호와 유용성 간의 상충 관계를 보이는 것을 확인했습니다. 재학습 및 미세 조정은 종종 균일한 GA+Refit 방식보다 더 강력한 유용성 유지 성능을 제공하는 반면, TrajDeleter는 여전히 유용한 비교 대상이지만 동일한 감사 기준 하에서 항상 더 우수한 성능을 보이지는 않습니다. 또한, 참조 모델 기반 공격, 임계값 공격, 편차 공격, 동등성 공격, 액션 오류 공격, 표현 기반 공격, 그리고 쿼리 제한 공격 등을 통해 단일 확률 기반 멤버십 점수가 삭제 품질을 과장할 수 있다는 것을 확인했습니다. 따라서 평가된 환경에서 오프라인 RL 데이터 삭제에 대한 결론은 단일 점수 감사만으로는 안정적이지 않으며, 매칭된 비멤버 그룹 구성, 재학습 기준의 보정, 공격 유형, 유지된 유용성, 그리고 진단 아키텍처 또는 구성 요소 수준의 증거를 포함하는 다양한 요인에 따라 달라집니다.

Original Abstract

Offline Reinforcement Learning (RL) agents are trained on fixed behavioral trajectories, which makes trajectory-level deletion important when selected data must be removed after training. Evaluating such deletion is difficult because a lower membership score can reflect trajectory removal, residual memorization visible to another attack, or policy collapse that destroys useful behavior. We introduce Trajectory-level memOrization and Unlearning in offline RL (TOUR), a benchmark that combines trajectory-level partitioning, matched non-member controls, retraining references, retained-performance anchors, and multi-attack privacy auditing. Across D4RL locomotion experiments and an exploratory AntMaze extension, TOUR shows that common deletion baselines have environment-dependent privacy-utility behavior. Retraining and fine-tuning often provide stronger retained-utility references than uniform GA+Refit, while TrajDeleter remains a useful comparator but is not uniformly stronger under the same audit. Reference-model, threshold, deviation, equivalence, action-error, representation-based, and query-limited attacks further show that a single likelihood-based membership score can overstate deletion quality. In the evaluated settings, conclusions about offline RL unlearning are therefore not stable under single-score auditing. They depend on matched non-member construction, retraining-relative calibration, attack family, retained utility, and explicit scope for diagnostic architecture or component-level evidence.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!