2606.16952v1 Jun 15, 2026 cs.LG

환상과 노출: 합성 데이터 감사에 대한 인과적 프레임워크

Phantoms and Disclosures: a Causal Framework for Auditing Synthetic Data

Adel Javanmard
Adel Javanmard
Citations: 57
h-index: 5
Sergei Vassilvitskii
Sergei Vassilvitskii
Citations: 3,380
h-index: 6
Kareem Amin
Kareem Amin
Citations: 74
h-index: 5
Rudrajit Das
Rudrajit Das
Citations: 362
h-index: 9
Alessandro Epasto
Alessandro Epasto
Citations: 3,283
h-index: 4
Dennis Kraft
Dennis Kraft
Citations: 188
h-index: 7
Mónica Ribero
Mónica Ribero
Citations: 15
h-index: 2

생성형 AI 및 대규모 언어 모델(LLM)의 빠른 발전은 민감한 실제 데이터 세트에 대한 개인 정보 보호 대안으로 합성 데이터에 대한 관심을 불러일으켰습니다. 그러나 고유용성을 가진 합성 데이터를 생성하는 것은 종종 훈련 데이터에서 개인 정보를 기억하고 재현할 위험을 수반합니다. 이 연구에서는 이러한 데이터 노출을 탐지하고 설명하도록 설계된 사용자 정의 가능한 경험적 감사 프레임워크를 제시합니다. 당사의 프레임워크는 시스템이 사용자의 정보를 직접 복제하는 "실제 노출"과 시스템이 우연히 사용자의 데이터를 생성하는 "환상 노출"을 구별하는 메커니즘을 도입합니다. 입력 데이터를 훈련 세트와 검증 세트로 분할하고 엄격한 통계적 가설 검정을 적용하여 관찰된 노출이 엄격한 개인 정보 보호 기준(예: 제로 러닝 또는 특정 차등 프라이버시(DP) 경계)과 일치하는지 여부를 판단합니다. 중요한 점은 이 접근 방식은 모델 액세스, 캐나리 삽입 또는 참조 모델 훈련을 필요로 하지 않으며, 오직 합성 출력과 검증 세트만 필요합니다. 당사는 이 프레임워크가 효과적으로 멤버십 추론 공격으로 기능하며, 기존의 데이터 기반 감사 방법보다 더 엄격한 개인 정보 유출에 대한 경험적 하한을 제공한다는 것을 보여줍니다. 당사의 접근 방식은 모델에 독립적이며 모든 합성 데이터 생성 메커니즘에 적용 가능하며, 섀도 모델 또는 캐나리 기반 대안보다 훨씬 적은 계산 자원을 필요로 합니다.

Original Abstract

The rapid adoption of generative AI and Large Language Models (LLMs) has spurred interest in synthetic data as a privacy-preserving alternative to sensitive real-world datasets. However, generating high-utility synthetic data often carries the risk of memorizing and regurgitating private information from the training corpus. In this work, we present a customizable empirical auditing framework designed to detect and explain such data disclosures. Our framework introduces a mechanism to distinguish between "true disclosures"-where the system directly reproduces a user's information-and "phantom disclosures''-where the system incidentally generates a user's data. By partitioning input data into training and holdout sets and applying rigorous statistical hypothesis testing, we determine if observed disclosures are consistent with strict privacy baselines, such as zero-learning or specific Differential Privacy (DP) bounds. Crucially, this approach requires no model access, no canary insertion, and no reference model training -only the synthetic output and a held-out control set. We demonstrate that this framework effectively functions as a membership inference attack, providing empirical lower bounds on privacy leakage that are tighter than prior data-based auditing methods. Our approach is model-agnostic, applies to any synthetic data generation mechanism, and requires orders of magnitude fewer computational resources than shadow-model or canary-based alternatives.

0 Citations
0 Influential
4.5 Altmetric
22.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!