2608.04519v1 Aug 05, 2026 cs.AI

누수 방지 학습 제거: 다중 단계 추론 일관성 및 복구 강건성을 평가하기 위한 새로운 벤치마크

Leak-Resistant Unlearning: A New Benchmark for Evaluating Multi-Hop Reasoning Consistency and Recovery Robustness

Han Qiu
Han Qiu
Citations: 19
h-index: 3
Zhicong Huang
Zhicong Huang
Citations: 186
h-index: 6
Cheng Hong
Cheng Hong
Citations: 34
h-index: 4
Haoting Qian
Haoting Qian
Citations: 50
h-index: 3
Qingjie Zhang
Qingjie Zhang
Citations: 166
h-index: 5

머신 러닝 모델에서 민감한 정보가 제거되었는지 여부를 이해하는 데 있어 학습 제거 방법의 성능을 측정하는 것은 매우 중요합니다. 현재의 학습 제거 벤치마크는 주로 단일 단계 질문과 제한된 수의 다중 단계 질문을 포함하고 있습니다. 이러한 벤치마크는 효과적이지만, 다음과 같은 두 가지 문제점에 직면해 있습니다. (1) 지식은 고립되어 있지 않으므로, 다양한 다중 단계 추론 경로가 일반적인 쿼리보다 더 큰 정보 누출을 유발할 수 있습니다. (2) 학습 제거는 취약할 수 있으며, 경량화된 후처리 적응과 같은 복구 공격을 통해 학습 제거된 지식이 부분적으로 복구될 수 있으므로, 정적인 평가는 충분하지 않습니다. 따라서 본 논문에서는 다양한 추론 경로와 복구 공격에 대한 강력한 LLM(Large Language Model) 지식 제거를 이해하기 위한 새로운 벤치마크인 { { }을 소개합니다. 우리는 이 벤치마크를 사용하여 3개의 모델, 6가지 학습 제거 방법 및 2개의 신중하게 선별된 데이터 세트를 실험했습니다. 결과는 기존 방법이 다중 단계 추론 경로와 복구 공격에 취약하다는 것을 보여줍니다. 또한 LLM 학습 제거에서 삭제 성능, 강건성 및 모델 유용성 간의 균형을 추가적으로 탐색합니다.

Original Abstract

Benchmarking machine unlearning methods is critical to understand whether sensitive knowledge is removed from large language models (LLMs) or not. Current unlearning benchmarks include mainly single-hop questions and a narrow set of multi-hop questions. Although effective, they still face two challenges. (1) Knowledge is not isolated, whereby diverse multi-hop reasoning paths can potentially induce knowledge leakage than normal queries. (2) Unlearning may be fragile: unlearned knowledge can be partially recovered through recovery attacks such as lightweight post-unlearning adaptation, making static evaluation insufficient. Therefore, in this paper, we introduce \unlearning as a novel benchmark to understand robust LLM knowledge removal across diverse reasoning paths and recovery attacks. We experiment with this benchmark on 3 models, 6 unlearning methods, and 2 carefully curated datasets. Results show that existing methods are vulnerable to multi-hop reasoning paths and recovery attacks. We further explore the trade-off among forget quality, robustness, and model utility for LLM unlearning.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!