2606.31664v1 Jun 30, 2026 cs.CV

생체 인식 검증을 위한 희소성 유도 다이버전스 손실

Sparsity-Inducing Divergence Losses for Biometric Verification

Themos Stafylakis
Themos Stafylakis
Omilia Conversational Intelligence
Citations: 3,615
h-index: 30
Dimitrios Koutsianos
Dimitrios Koutsianos
Citations: 8
h-index: 1
Ladislav Movsner
Ladislav Movsner
Citations: 91
h-index: 6
Yannis Panagakis
Yannis Panagakis
Citations: 3,579
h-index: 30

얼굴 및 음성 인식 검증의 성능은 CosFace 및 ArcFace와 같은 마진 페널티 소프트맥스 손실에 크게 의존합니다. 최근에 소개된 α-다이버전스 손실 함수는 특히 α > 1일 때 희소한 해를 유도하는 능력 덕분에 매력적인 대안을 제시합니다. 그러나 표준 기하학적 마진은 소프트맥스 함수에 맞춰 설계되었으며, 이 일반화된 확률 프레임워크로 자연스럽게 확장되지 않습니다. 본 논문에서는 새로운 α-다이버전스 손실인 Q-Margin을 제안합니다. Q-Margin은 원칙적인 확률 기반 마진을 도입합니다. 기존 방법과는 달리, Q-Margin은 기하학적 페널티를 로짓(정규화되지 않은 로그 우도)에 적용하는 대신, 마진 페널티를 참조 측정값(사전 확률)에 직접 인코딩합니다. 이러한 제형은 자연스럽게 판별력 있는 임베딩을 장려하면서 α-다이버전스의 유용한 희소성 특성을 유지합니다. 우리는 Q-Margin이 어려운 IJB-B 및 IJB-C 얼굴 인식 검증 벤치마크에서 경쟁력 있거나 우수한 성능을 달성하며, VoxCeleb의 음성 인식에서도 유사하게 강력한 결과를 보인다는 것을 입증했습니다. 특히 동일한 학습 레시피로 훈련된 ArcFace 및 CosFace 기준 모델과 비교했을 때, Q-Margin은 낮은 오탐율(FAR)에서 지속적으로 개선되는 성능을 보여주며, 이는 실제 고보안 애플리케이션에 매우 중요합니다. 마지막으로, Q-Margin의 극단적인 희소성은 정확하고 메모리 효율적인 학습을 가능하게 하며, 수백만 개의 아이덴티티를 가진 데이터 세트에 대한 확장 가능한 솔루션을 제공합니다.

Original Abstract

Performance in face and speaker verification is largely driven by margin-penalty softmax losses such as CosFace and ArcFace. Recently introduced $α$-divergence loss functions offer a compelling alternative, particularly due to their ability to induce sparse solutions (when $α>1$). However, standard geometric margins are designed for the softmax function and do not naturally extend to this generalized probabilistic framework. In this paper we propose Q-Margin, a novel $α$-divergence loss that introduces a principled probabilistic margin. Unlike conventional methods that apply geometric penalties to the logits (unnormalized log-likelihoods), Q-Margin encodes the margin penalty directly into the reference measure (prior probabilities). This formulation naturally encourages discriminative embeddings while preserving the beneficial sparsity properties of the $α$-divergence. We demonstrate that Q-Margin achieves competitive or superior performance on the challenging IJB-B and IJB-C face verification benchmarks and similarly strong results in speaker verification on VoxCeleb. Crucially, against ArcFace and CosFace baselines trained under an identical recipe, Q-Margin consistently improves at low False Acceptance Rates (FARs), a capability critical for practical high-security applications. Finally, the extreme sparsity of the Q-Margin posteriors enables exact and memory-efficient training, offering a scalable solution for datasets with millions of identities.

0 Citations
0 Influential
15 Altmetric
75.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!