2607.27968v1 Jul 30, 2026 cs.LG

이분법적 보상을 넘어서: 강화 학습 기반 지식 제거(Reinforcement Unlearning)를 위한 보상 설계 비교 연구

Beyond Binary Rewards: A Comparative Study of Reward Design for Reinforcement Unlearning

Bardh Prenkaj
Bardh Prenkaj
Citations: 3
h-index: 1
Gjergji Kasneci
Gjergji Kasneci
Citations: 17,895
h-index: 40
Efstratios Zaradoukas
Efstratios Zaradoukas
Citations: 6
h-index: 2
Davide Gabrielli
Davide Gabrielli
Citations: 13
h-index: 1

머신 러닝에서 지식 제거는, GDPR 및 EU AI Act과 같은 개인 정보 보호 규정에 따라, 전체 재학습 없이 훈련된 언어 모델로부터 특정 지식을 선택적으로 제거하는 것을 목표로 합니다. 최근 연구에서는 지식 제거를 검증 가능한 보상을 기반으로 하는 강화 학습(RLVR) 문제로 재정의했습니다. 여기서 모델은 직접 출력으로부터 계산되는 검증 가능한 보상에 맞춰 최적화됩니다. 그러나 기존 방법은 제한적인 학습 신호를 제공하는 희소한 이분법적 보상에 의존하며, 이는 금지된 콘텐츠가 회피되었는지 여부만을 나타내어 수렴 속도를 제한합니다. 본 논문에서는 강화 지식 제거(RUL) 프레임워크 내에서 보상 설계가 지식 제거 효율성에 미치는 영향을 연구합니다. 우리는 검증 가능성과 희소성을 분리하는 원칙적인 보상 분해 프레임워크를 도입하고, 두 가지 새로운 보상 함수를 제안합니다. 첫째는 금지된 개념의 발생 횟수에 따라 단계별 벌점을 제공하는 지수형 보상이며, 둘째는 PageRank에서 영감을 받아 의미적 중요도에 따라 벌점을 가중치는 페이지랭크 기반 보상입니다. 우리는 Real World Knowledge Unlearning (RWKU) 벤치마크를 사용하여 실험을 수행했습니다. 그 결과, 제안된 두 가지 보상 함수 모두 이분법적 설정보다 일관되게 우수한 성능을 나타냈으며, 유사한 망각 성능을 $3 imes$ 더 빠르게 달성하고 일반적인 모델 유용성을 유지했습니다. 이러한 결과는 보상 설계가 지식 제거 효율성의 핵심 요소이며, 확장 가능하고 효율적인 머신 러닝 지식 제거를 위한 실질적인 경로를 제시한다는 것을 보여줍니다.

Original Abstract

Machine unlearning seeks to selectively remove specific knowledge from trained language models without full retraining, a growing necessity under privacy regulations such as GDPR and the EU AI Act. Recent work has reformulated unlearning as a Reinforcement Learning with Verifiable Rewards (RLVR) problem, where models are optimized against verifiable rewards computed directly from their outputs. However, existing methods rely on sparse binary rewards that provide minimal learning signal, indicating only whether forbidden content was avoided, and limiting convergence speed. In this paper, we study how reward design affects unlearning efficiency within the Reinforcement Unlearning (RUL) framework. We introduce a principled reward decomposition framework that decouples verifiability from sparsity, and propose two new reward functions: an exponential reward that provides graded penalties based on the count of forbidden-concept occurrences, and a PageRank inspired reward that weights penalties by semantic importance. We conduct experiments on the Real World Knowledge Unlearning (RWKU) benchmark, demonstrating that both rewards consistently outperform the binary setting, while reaching similar forgetting performance up to $3\times$ faster and preserving general model utility. Our results show that reward design is a key driver of unlearning efficiency offering a practical path toward scalable and efficient machine unlearning.

0 Citations
0 Influential
20 Altmetric
100.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!