2606.06286v1 Jun 04, 2026 cs.CL

LLM은 학습 데이터를 유출할 수 있지만, 정말로 그렇게 하고 싶어하는가? LLM의 기억 능력에 대한 경향성 기반 평가

LLMs Can Leak Training Data But Do They Want To? A Propensity-Aware Evaluation of Memorization in LLMs

Peter Schneider-Kamp
Peter Schneider-Kamp
Citations: 174
h-index: 6
Lukas Galke Poech
Lukas Galke Poech
Citations: 5
h-index: 1
Gianluca Barmina
Gianluca Barmina
Citations: 18
h-index: 2

대규모 언어 모델(LLM)은 학습 데이터를 재현할 수 있지만, 기존의 기억 능력 평가는 주로 모델이 그러한 행동을 하도록 강제될 수 있는지 여부를 측정하는 데 집중합니다. 본 연구에서는 접두사 기반 공격과 비적대적인 평가를 비교하는 경향성 기반 기억 능력 평가 프레임워크인 PropMe를 소개합니다. 제안된 메트릭 변환은 기존 함수에 적용하여 경향성을 측정할 수 있도록 합니다. 또한, infini-gram을 기반으로 구축된 경량 추적 파이프라인인 SimpleTrace를 통해 모델의 생성 결과를 대규모 학습 데이터 코퍼스와 결정적으로 연결하고, 완전한 재현, 유사한 재현, 그리고 경향성 변환된 기억 능력 메트릭을 계산합니다. 두 개의 완전히 공개된 모델(Comma 및 DFM Decoder)을 사용하여 Common Pile과 Dynaword 두 가지 데이터셋에 대해 한국어와 영어로 평가한 결과, 공격적인 시나리오와 일반적인 사용 환경 간에 일관된 차이가 나타났습니다. 접두사 기반 공격은 일반적이거나 데이터셋 특화된 프롬프트보다 훨씬 강력한 기억 능력 신호를 유발하는 반면, 전반적으로 경향성 점수는 낮게 유지되었습니다. 따라서 LLM은 직접적인 자극을 받을 때 학습 데이터를 드러낼 수 있지만, 일반적인 비적대적인 환경에서는 거의 그러지 않습니다. 또한, Comma에서 지속적으로 사전 훈련된 DFM Decoder는 Common Pile에 대한 기억 능력과 기억 유출 가능성이 감소하는 경향을 보였으며, 이는 이후의 훈련이 부분적으로 다른 데이터에 초점을 맞추면 기억 능력이 감소할 수 있음을 시사합니다. 본 연구 결과는 기억 능력 평가 보고서가 최악의 경우 추출 가능성과 일반적인 유출 가능성을 모두 포함해야 한다는 점을 강조하며, 이를 통해 이 현상에 대한 보다 포괄적인 이해를 제공하고자 합니다.

Original Abstract

Large language models can reproduce training data, but existing memorization evaluations mostly measure whether models can be forced to do so, rather than whether they do so under ordinary use. We introduce PropMe, a propensity-aware framework for memorization evaluation that contrasts prefix-based capability attacks with non-adversarial evaluations. We propose a metric transformation that, applied to existing functions, allows to create propensity metrics. We further introduce SimpleTrace, a lightweight tracing pipeline built on infini-gram that deterministically attributes model generations to large-scale training corpora and computes verbatim, near-verbatim, and propensity-transformed memorization metrics. Evaluating two fully-open models: Comma and DFM Decoder on two datasets: Common Pile and Dynaword in two languages, we find a consistent gap between capability and propensity: prefix attacks elicit substantially stronger memorization signals than generic or dataset-specific prompts, while propensity scores remain low overall. Thus, the models can reveal training data when directly elicited, but rarely do so in more common non-adversarial settings. We also find that DFM Decoder, which is continually pre-trained from Comma, exhibits reduced memorization and memorization propensity for Common Pile, confirming that memorization capability can decrease when later training emphasizes partially different data. Our results suggest, and we encourage, that memorization audits should report both worst-case extractability and ordinary leakage propensity in order to have a more comprehensive view of this phenomenon.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!