불확실성 정량화를 통한 실존/가짜 분류를 위한 해석 가능한, 인간 레이블이 없는 심층 학습
Interpretable Human-Label-Free Deep Learning for Real-Bogus Classification with Uncertainty Quantification
시간 영역 관측 데이터는 많은 일시적인 후보를 생성하며, 이는 자동화된 발견 파이프라인에서 중요한 단계인 실존/가짜 분류로 이어집니다. 신뢰할 수 있는 레이블은 비용이 많이 들고, 커뮤니티 레이블은 노이즈가 많고 관측 데이터에 의존적입니다. 본 연구에서는 인간 레이블이 없는 데이터만을 사용하여 훈련 가능하고, 심각한 클래스 오염에도 강건하며, 보정된 불확실성 정량화를 제공하는 실존/가짜 분류 프레임워크를 개발하는 것을 목표로 합니다. 시뮬레이션된 일시적인 신호 주입과 오염된 관측 데이터 클래스를 결합하여, 비대칭 코-티칭을 사용하여 서로 다른 레이블 노이즈 수준을 가진 모델을 훈련합니다. 표준 벤치마크 데이터 세트에서 성능을 평가하고, 잠재 공간 시각화 도구를 사용하여 학습된 표현을 분석합니다. 불확실성 정량화를 위해 MC 드롭아웃과 딥 앙상블을 비교하고, 이중 네트워크 설정을 활용하여 보정을 개선하는 저비용 하이브리드 전략을 제안합니다. 평가를 광도 곡선 영역으로 확장하여 광도 곡선 클래스의 복원력을 평가합니다. 본 방법은 레이블이 지정된 데이터 세트에서 뛰어난 실존/가짜 성능을 달성하며, 심각한 클래스 오염에도 안정적입니다. 또한, 일시적인 광도 곡선 클래스를 높은 정확도로 복원하지만, 단일 소스 식별은 광도 곡선에서 파생된 레이블의 모호성으로 인해 제한됩니다. 제안하는 하이브리드 불확실성 정량화 방법은 더 비용이 많이 드는 앙상블 기반 모델과 비교하여 경쟁력 있는 보정 성능을 제공합니다. 잠재 공간 분석 결과, 불확실성이 의사 결정 경계와 일치하며, 가짜 데이터 집단 내의 하위 클래스를 드러냅니다. 본 연구 결과는 주입 기반의 약한 감독 학습이 인간 레이블이 없는 훈련 데이터를 사용하지 않고도 확장 가능하고 일관된 실존/가짜 분류를 가능하게 하며, 동시에 보정된 불확실성을 제공할 수 있음을 보여줍니다. 또한, 이 방법은 향후 관측 데이터에 적용하기 위해 주입 기반의 훈련 파이프라인을 재실행하여 쉽게 이전할 수 있습니다.
Time-domain surveys generate many transient candidates, making Real-Bogus classification a critical step in automated discovery pipelines. Reliable labels are costly, while community labels can be noisy and survey-dependent. We aim to develop a Real-Bogus classification framework that can be trained without human-labeled data using injected transients and bogus-dominated survey data, remains robust under strong class contamination, and provides calibrated uncertainty quantification. We combine simulated transient injections with a contaminated survey class and train a dual-network model using asymmetric co-teaching for classes with different label-noise levels. We evaluate performance on a benchmark subset and analyze the learned representation with latent-space visualization tools. For uncertainty quantification (UQ), we compare MC dropout and deep ensembles and propose a low-cost hybrid strategy that exploits the dual-network setting to improve calibration. We extend the evaluation to the light-curve domain to assess recovery of light-curve classes. The method achieves strong Real-Bogus performance on the labeled subset and remains stable under severe class contamination. It recovers transient light-curve classes with high fidelity, while single-source identification is limited by ambiguity in light-curve-derived labels. Our hybrid UQ approach achieves competitive calibration relative to more expensive ensemble baselines. Latent-space analyses indicate that uncertainty aligns with the decision boundary and reveal subclasses within the bogus population. Our results show that injection-driven, weakly supervised training can enable scalable and consistent Real-Bogus classification without human-labeled training data while providing calibrated uncertainties. The method is suited for transfer to forthcoming surveys by re-running the injection-based training pipeline.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.