사후 심플렉스 기하를 이용한 라벨 없는 다중 분류
Multiclass Classification without Labels via Posterior Simplex Geometry
많은 분류 문제에서 신뢰할 수 있는 개별 데이터 레벨의 라벨이 제공되지 않는 경우가 많습니다. 하지만, 종종 약하게 보강된 비라벨 샘플을 구성하는 것이 가능합니다. 이러한 샘플은 서로 다른 기준, 출처, 모집단 또는 실험 조건에 의해 선택되며, 이는 숨겨진 클래스 비율을 변경하지만 이를 드러내지는 않습니다. 라벨 없는 분류 (CWoLa)는 이진 경우 ($K=2$)에서, 서로 다른 클래스 비율을 가진 두 개의 불순 혼합물을 구별하도록 훈련된 분류기는 혼합 비율을 알지 못하는 상태에서도 최적의 클래스 구분기를 복원할 수 있음을 보여줍니다. 우리는 이러한 원리를 여러 개의 비라벨 혼합물 ($K>2$)로부터 다중 클래스 학습으로 확장합니다. 여기서 학습자는 혼합물의 종류만 관찰하며, 숨겨진 클래스 라벨이나 클래스 사전 행렬은 알지 못합니다. 우리는 다중 클래스 혼합 모델에 대해 베이즈 최적 혼합 분류기 $g^ riangle$가 데이터 포인트를 혼합-사후 공간에 내장된 $(K-1)$-심플렉스로 매핑한다는 것을 증명했습니다. 이 심플렉스의 $K$개의 꼭짓점은 숨겨진 클래스를 통해 알려지지 않은 혼합 행렬에 의해 유도됩니다. 이러한 기하학적 구조를 활용하여, 우리는 사전 정보 없이 표준 분류기를 훈련시켜 혼합물의 종류를 구별하도록 하고, 그 후 사후 심플렉스 적합 또는 병목(bottleneck) 아키텍처를 사용하여 숨겨진 클래스 구조를 추출하는 절차를 제안합니다. MNIST, CIFAR-10 및 Galaxy10 DECaLS 데이터에 대한 실험 결과는 혼합물의 종류만으로도 숨겨진 클래스와 혼합물 내에서의 비율을 복원할 수 있음을 보여줍니다. 우리는 약하게 감독된 방법과 완전하게 감독된 방법의 성능 격차를 줄여, 라벨이 부족한 영역에서 다중 클래스 발견을 위한 수학적으로 정당하고 확장 가능한 도구를 제공합니다.
In many classification problems, reliable instance-level labels are unavailable. However, it is often possible to construct weakly enriched unlabeled samples: datasets selected by different cuts, sources, populations, or experimental conditions that change latent class proportions without revealing them. Classification without Labels (CWoLa) shows that, in the binary case ($K=2$), a classifier trained to distinguish two impure mixtures with different class proportions can recover an optimal class discriminator without knowing the mixture proportions. We extend this principle to multiclass learning from several unlabeled mixtures ($K>2$), where the learner observes only mixture identity and neither latent class labels nor class-prior matrices. We prove that, for a multiclass mixture model, the Bayes-optimal mixture classifier $g^\star$ maps data points into a $(K-1)$-simplex embedded in mixture-posterior space. The $K$ vertices of this simplex are induced by the latent classes through the unknown mixing matrix. Leveraging this geometry, we propose prior-free procedures that train a standard classifier to distinguish mixture identities and then extract latent class structure using either post-hoc simplex fitting or a bottleneck architecture. Experiments on MNIST, CIFAR-10, and Galaxy10 DECaLS show that mixture identity alone can recover latent classes and their fractions in the mixture. By narrowing the gap between weakly supervised and fully supervised performance, we provide a mathematically grounded, scalable tool for multiclass discovery in label-scarce domains.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.