2608.10462v1 Aug 11, 2026 cs.CL

LLM 데이터 오염 탐지를 위한 사전 학습 후 특징 변화 보정

Calibrating Post-Training Feature Shifts for LLM Data Contamination Detection

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Jianwei Wang
Jianwei Wang
Citations: 116
h-index: 5
Wenjie Zhang
Wenjie Zhang
Citations: 37
h-index: 4
Mozhu Zhou
Mozhu Zhou
Citations: 3
h-index: 1
Mengqi Wang
Mengqi Wang
Citations: 14
h-index: 2
Zheng Yang
Zheng Yang
Citations: 0
h-index: 0

대규모 언어 모델(LLM)은 방대한 양의 데이터로 학습되는데, 이 데이터에는 저작권이 있거나 개인정보를 포함할 수 있는 내용이 있을 수 있습니다. 따라서 데이터 오염 탐지(DCD)는 특정 텍스트가 대상 LLM의 사전 학습 코퍼스에 속하는지를 판단하는 것을 목표로 합니다. 최근의 최첨단 DCD 방법은 입력 텍스트와 해당 모델 출력을 기반으로 멤버십 특징을 추출하는 특징 기반 패러다임을 따릅니다. 그러나 대부분의 현대적인 LLM은 명령어 조정, 선호도 최적화 및 추론 중심 학습과 같은 사후 학습 과정을 거치는데, 이는 모델 출력에 변화를 주고 해당 멤버십 특징을 이동시켜 멤버와 비멤버 간의 구별력을 감소시킬 수 있습니다. 이 문제를 해결하기 위해, 우리는 널리 적용 가능한 특징 기반 DCD 방법의 보정 프레임워크인 CalibDCD를 제안합니다. 이는 (1) 멀티뷰 변화 감지(Multi-View Shift Detection), 즉 사후 학습과 관련된 반복적인 특징 변화를 식별하고, (2) 경계가 있는 특징 보정(Bounded Feature Correction), 즉 이러한 변화의 영향력을 선택적으로 완화하여 멤버십 예측에 미치는 영향을 줄이는 것을 포함합니다. 구체적으로, 멀티뷰 변화 감지는 알려진 비멤버 텍스트에 대한 제어된 프롬프트 변형을 평가하고 가장 유용한 정보를 제공하는 관점을 통합하여 반복적인 특징 변화를 식별합니다. 경계가 있는 특징 보정은 검출된 변화와 일치하는 특징 요소를 선택적으로 조정하고, 유용한 탐지 정보를 유지하기 위해 보정 정도를 제어합니다. 실험 결과는 CalibDCD가 기존의 특징 기반 감지기 성능을 꾸준히 향상시키며, AUC에서 최대 7.0%, TPR@5%FPR에서 15.0%의 성능 향상을 달성함을 보여줍니다.

Original Abstract

Large language models (LLMs) are trained on massive and largely undisclosed corpora that may contain copyrighted or privacy-sensitive content. Data contamination detection (DCD) therefore aims to determine whether a given text is a member of the pre-training corpus of a target LLM. Recent state-of-the-art DCD methods follow a feature-based paradigm that derives membership features from the input text and the corresponding model output. However, most modern LLMs undergo post-training, such as instruction tuning, preference optimization, and reasoning-oriented training, which can alter model outputs and shift the corresponding membership features, thereby reducing the separability between members and non-members. To address this problem, we propose CalibDCD, a broadly applicable calibration framework for feature-based DCD methods, comprising (1) Multi-View Shift Detection, which identifies recurring feature shifts associated with post-training, and (2) Bounded Feature Correction, which selectively mitigates their influence on membership prediction. Specifically, Multi-View Shift Detection evaluates controlled prompt variants on known non-member texts and consolidates the most informative views to identify recurring feature shifts. Bounded Feature Correction selectively adjusts feature components aligned with the detected shifts and controls the correction extent to preserve useful detection information. Experiments show that CalibDCD consistently improves existing feature-based detectors, with gains of up to 7.0% in AUC and 15.0% in TPR@5%FPR.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!