인공 합성 얼굴 데이터셋 식별을 위한 행동 특징 임베딩
Have I Seen You? Embedding Behavior Signals Synthetic Face Dataset Membership
합성 얼굴 데이터셋은 생체 인식 분야에서 개인 정보 노출 및 데이터 접근 제한을 줄이는 데 점점 더 많이 사용되고 있습니다. 그러나 이러한 데이터셋을 생성하는 모델들은 실제 얼굴 데이터를 기반으로 학습되기 때문에, 합성 데이터는 여전히 실제 원본 데이터에 대한 정보를 드러낼 수 있습니다. 본 연구에서는 데이터셋 수준의 멤버십 추론 공격을 통해 이 위험성을 분석합니다. 먼저 얼굴 인식 모델을 훈련하는 데 사용된 합성 데이터셋을 식별하고, 그 다음에는 해당 생성 모델을 훈련하는 데 사용된 실제 데이터셋을 추론합니다. 11개의 얼굴 인식 모델, 11개의 합성 데이터셋, 그리고 7개의 실제 데이터셋에 대한 실험 결과, 공격은 100%의 정확도로 합성 학습 데이터셋을 복구하고, 생성 모델의 원본 데이터셋을 54.5%의 정확도로 식별했습니다. 이러한 결과는 합성 데이터가 여전히 실제 학습 데이터셋의 정보를 포함할 수 있으며, 개인 정보 보호를 위한 시스템 구축에는 더 강력한 정보 유출 방지 기술이 필요하다는 것을 보여줍니다.
Synthetic face datasets are increasingly used to reduce privacy exposure and data access constraints in biometric recognition. Yet the generators that produce these datasets are trained on real faces, so synthetic data may still reveal their real source data. We study this risk through a dataset-level membership inference attack that first identifies the synthetic dataset used to train a face recognizer and then infers the real dataset used to train the generator. Across 11 face recognition models, 11 synthetic datasets, and 7 real datasets, the attack recovers the synthetic training dataset in 100% of cases and identifies the generator's source dataset in 54.5% of cases. These results show that synthetic data can retain dataset-level traces of real training data and that privacy-preserving deployment requires stronger leakage mitigation.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.