연속적인 LLM 지식 삭제를 위한 경로 기반 망각-복구 네트워크
Trajectory-Guided Forget-Recover Network for Continual LLM Unlearning
머신 러닝 모델에서 데이터 삭제는 민감한 정보의 영향을 제거하는 것을 목표로 합니다. 현실 세계에서는 이러한 데이터 삭제 요청이 지속적으로 발생하며, 이로 인해 두 가지 어려움이 발생합니다. 첫째, 데이터 삭제 과정에서 관련 계산량이 남아있는 경로로 재분배되어, 이전에는 삭제된 지식이 다시 나타날 수 있습니다. 둘째, 반복적인 데이터 삭제는 모델의 유용성을 유지하는 데 필요한 용량을 점진적으로 감소시킬 수 있습니다. 이러한 문제를 해결하기 위해, 본 연구에서는 경로 기반 망각-복구 네트워크(Trajectory-guided Forget-Recover Network, TFR-Net)를 제안합니다. TFR-Net은 각 요청에 따른 채널 수준의 위험을 추적하며, 지속적인 관련 채널과 일시적인 활성 영역을 분리하여 지속적인 채널만을 삭제합니다. 또한, TFR-Net은 잠재되어 있던 채널을 재활성화하여 모델 용량을 복구합니다. 이 채널들은 유용성 유지에 큰 기여를 하며, 현재 및 과거의 망각 위험이 낮습니다. 복구는 유지되는 유용성의 저하가 미리 정의된 허용 범위 내에 있을 때만 수락됩니다. 네 가지 데이터 세트에 대한 실험 결과, TFR-Net은 대표적인 기본 모델보다 삭제 효과와 유지되는 유용성 간의 균형을 더 나은 방식으로 달성하는 것을 보여줍니다.
Machine unlearning aims to eliminate the influence of sensitive data on a model. In the real world, unlearning requests arrive continually, which gives rise to two challenges. First, an unlearning intervention may redistribute target-related computation across remaining pathways, allowing previously forgotten knowledge to re-emerge. Second, repeated unlearning interventions may progressively reduce the model capacity needed to preserve retained utility. To address these challenges, we propose the Trajectory-guided Forget-Recover Network (TFR-Net). TFR-Net tracks channel-level risk across requests. It separates persistent target-related channels from transient hotspots and suppresses only the persistent ones. TFR-Net also recovers model capacity by reactivating dormant channels. These channels make strong contributions to retained utility and show low current and historical forget risk. The recovery is accepted only when retained-utility degradation remains within a predefined tolerance. Experiments on four datasets show that TFR-Net consistently achieves a more favorable trade-off between unlearning effectiveness and retained utility than representative baselines.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.