2607.27764v1 Jul 30, 2026 cs.CV

아이덴티티 분리 및 기하학적 특성 보존을 통한 개인 얼굴 인식 학습 데이터셋 공개 방법

Private Face Recognition Training Dataset Publication via Identity-Decoupled and Geometry-Preserving Face Distillation

Haoyuan Zhang
Haoyuan Zhang
Citations: 32
h-index: 2
Xiangyu Zhu
Xiangyu Zhu
Citations: 7,984
h-index: 33
Siran Peng
Siran Peng
Citations: 90
h-index: 4
Tianshuo Zhang
Tianshuo Zhang
Citations: 261
h-index: 6
Shuhuan Chen
Shuhuan Chen
Citations: 11
h-index: 2
Weisong Zhao
Weisong Zhao
Citations: 112
h-index: 7
Xiao-Yu Zhang
Xiao-Yu Zhang
Citations: 238
h-index: 9
Haichao Shi
Haichao Shi
Citations: 23
h-index: 1
Zhen Lei
Zhen Lei
Citations: 5
h-index: 1

개인 얼굴 인식(FR) 학습 데이터셋을 공개하는 것은 프라이버시 문제와 관련이 있는데, 왜냐하면 얼굴은 개인 식별 정보를 담고 있기 때문이다. 개인 FR 학습 데이터셋의 프라이버시를 보호하기 위해, 실제 학습 데이터 대신 가짜 데이터를 공개하는 방법이 사용된다. 그러나 이러한 가짜 데이터를 사용하여 FR 모델을 훈련시키면 '아이덴티티 역설'이라는 문제가 발생한다: 즉, '가짜 얼굴이 인식 성능 향상에 유용한 정보(개인 식별 특징)를 포함하고 있기 때문에, 여전히 실제 개인과 연결될 수 있다.' 이상적인 가짜 얼굴은 원래의 아이덴티티와 분리되어야 하지만, 동시에 훈련을 위한 신뢰할 만한 아이덴티티 샘플 역할을 해야 한다. 이러한 특징들을 너무 과도하게 제거하면 인식 학습에 필요한 클래스 구조가 파괴될 수 있고, 반대로 너무 충실하게 보존하면 실제 개인과의 연결 가능성이 높아진다. 우리는 이 역설이 '원래 데이터의 아이덴티티 정보'와 '인식에 유용한 가짜 얼굴의 기하학적 특징'을 혼동하기 때문에 발생한다고 주장한다. 따라서, 원래 데이터와의 연관성을 줄이기 위해서는 원래 아이덴티티 정보를 억제해야 하고, 인식 학습을 위해서는 가짜 얼굴의 기하학적 특징을 보존해야 한다. 이러한 관점을 바탕으로, 우리는 '개인 얼굴 증류(Private Face Distillation)'라는 새로운 프레임워크를 제안한다. 이 방법은 원본 개인 정보로부터 분리된 가짜 아이덴티티를 생성하면서도, 초구면 기하학적 구조를 유지하는 '직교 기하학 보존' 기술과, 인식 학습을 위한 아이덴티티 관계를 유지하는 '관계 토폴로지 정렬' 기술을 사용한다. 다양한 환경에서 수행한 실험 결과, 개인 얼굴 증류 방법이 기존의 공개 데이터 방식을 능가하는 성능을 보여주었다. 예를 들어, IJB-C 데이터셋에서 'TAR@FAR=1e-3' 지표를 3.94% 향상시켰으며, 동시에 실제 개인과의 연결 가능성을 줄였다. 이러한 결과는 개인 FR 학습 데이터셋을 공개할 때, 원본 데이터와의 아이덴티티 연관성은 분리해야 하지만, 가짜 얼굴의 기하학적 특징은 보존해야 한다는 것을 시사한다.

Original Abstract

Publishing private face recognition~(FR) training datasets is privacy-sensitive because faces expose identity information. Private FR training dataset publication mitigates this risk by releasing protected proxies as substitutes for private training faces. However, training FR models with such data introduces an identity paradox: \emph{the identity cues that make released faces useful for recognition supervision are also the cues that make them linkable to real individuals.} A protected face should be decoupled from the original identity, yet still behave as a reliable identity sample for training. Removing these cues too aggressively may destroy the class structure needed for recognition learning, whereas preserving them too faithfully may increase source-identity linkability. We argue that this paradox stems from conflating source-aligned identity semantics with recognition-useful proxy identity geometry. The former should be suppressed to reduce linkage to private individuals, while the latter should be preserved for FR learning. Based on this insight, we propose \textbf{Private Face Distillation}, an identity-decoupling and geometry-preserving framework. It uses Orthogonal Geometry Preservation to construct decoupled proxy identities from private identity representations while maintaining hyperspherical geometry, and Relational Topology Alignment to preserve identity relations for recognition learning. Experiments across multiple domain-shifted FR scenarios show that Private Face Distillation achieves stronger utility than the evaluated publication baselines. On IJB-C surveillance, it improves $\mathrm{TAR}@\mathrm{FAR}{=}1\text{e-}{3}$ by 3.94\% over the baseline while reducing source-identity linkability. These results suggest that private FR training dataset publication should decouple source-identity correspondence while preserving proxy identity geometry.

0 Citations
0 Influential
16.5 Altmetric
82.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!