미세 조정된 언어 모델에서 의도치 않은 민감 정보 암기
Unintended Memorization of Sensitive Information in Fine-Tuned Language Models
대규모 언어 모델(LLM)을 민감한 데이터 세트에 대해 미세 조정하는 것은 개인 식별 정보(PII)의 의도치 않은 암기와 유출의 상당한 위험을 초래하며, 이는 개인 정보 보호 규정을 위반하고 개인의 안전을 위협할 수 있습니다. 본 연구에서는 모델 입력에만 나타나고 학습 대상에는 포함되지 않는 PII 노출이라는 중요한 취약점을 체계적으로 조사합니다. 합성 데이터와 실제 데이터를 사용하여, 의도치 않은 PII 암기를 정량화하기 위한 제어된 추출 방법을 설계하고, 언어, PII 빈도, 작업 유형, 모델 크기와 같은 요인이 암기 동작에 미치는 영향을 연구합니다. 또한, 차등 프라이버시, 머신 언러닝, 정규화, 선호도 정렬을 포함한 네 가지 개인 정보 보호 접근 방식을 벤치마킹하고, 개인 정보 보호와 작업 성능 간의 균형을 평가합니다. 결과는 일반적으로 사후 학습 방법이 더 일관된 개인 정보 보호-유용성 균형을 제공하지만, 특정 환경에서 차등 프라이버시가 유출 감소에 강력한 효과를 발휘한다는 것을 보여줍니다. 그러나 이는 학습 안정성을 저해할 수 있습니다. 이러한 결과는 미세 조정된 LLM에서 발생하는 암기 문제의 지속적인 어려움을 강조하며, 강력하고 확장 가능한 개인 정보 보호 기술의 필요성을 강조합니다.
Fine-tuning Large Language Models (LLMs) on sensitive datasets carries a substantial risk of unintended memorization and leakage of Personally Identifiable Information (PII), which can violate privacy regulations and compromise individual safety. In this work, we systematically investigate a critical and underexplored vulnerability: the exposure of PII that appears only in model inputs, not in training targets. Using both synthetic and real-world datasets, we design controlled extraction probes to quantify unintended PII memorization and study how factors such as language, PII frequency, task type, and model size influence memorization behavior. We further benchmark four privacy-preserving approaches including differential privacy, machine unlearning, regularization, and preference alignment, evaluating their trade-offs between privacy and task performance. Our results show that post-training methods generally provide more consistent privacy-utility trade-offs, while differential privacy achieves strong reduction in leakage in specific settings, although it can introduce training instability. These findings highlight the persistent challenge of memorization in fine-tuned LLMs and emphasize the need for robust, scalable privacy-preserving techniques.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.