2607.05393v1 Jul 06, 2026 astro-ph.IM

불확실성 정량화를 통한 실존/가짜 분류를 위한 해석 가능한, 인간 레이블이 없는 심층 학습

Interpretable Human-Label-Free Deep Learning for Real-Bogus Classification with Uncertainty Quantification

Raphaël Bonnet-Guerrini
Raphaël Bonnet-Guerrini
Citations: 4
h-index: 1
B. S'anchez
B. S'anchez
Citations: 23
h-index: 3
D. Fouchez
D. Fouchez
Citations: 18
h-index: 1
Benjamin Racine
Benjamin Racine
Citations: 18
h-index: 1
Maya Guy
Maya Guy
Citations: 0
h-index: 0
Mariam Sabalbal
Mariam Sabalbal
Citations: 0
h-index: 0
M. Yassine
M. Yassine
Citations: 2,346
h-index: 14
Vincenzo Piuri
Vincenzo Piuri
Citations: 175
h-index: 7

시간 영역 관측 데이터는 많은 일시적인 후보를 생성하며, 이는 자동화된 발견 파이프라인에서 중요한 단계인 실존/가짜 분류로 이어집니다. 신뢰할 수 있는 레이블은 비용이 많이 들고, 커뮤니티 레이블은 노이즈가 많고 관측 데이터에 의존적입니다. 본 연구에서는 인간 레이블이 없는 데이터만을 사용하여 훈련 가능하고, 심각한 클래스 오염에도 강건하며, 보정된 불확실성 정량화를 제공하는 실존/가짜 분류 프레임워크를 개발하는 것을 목표로 합니다. 시뮬레이션된 일시적인 신호 주입과 오염된 관측 데이터 클래스를 결합하여, 비대칭 코-티칭을 사용하여 서로 다른 레이블 노이즈 수준을 가진 모델을 훈련합니다. 표준 벤치마크 데이터 세트에서 성능을 평가하고, 잠재 공간 시각화 도구를 사용하여 학습된 표현을 분석합니다. 불확실성 정량화를 위해 MC 드롭아웃과 딥 앙상블을 비교하고, 이중 네트워크 설정을 활용하여 보정을 개선하는 저비용 하이브리드 전략을 제안합니다. 평가를 광도 곡선 영역으로 확장하여 광도 곡선 클래스의 복원력을 평가합니다. 본 방법은 레이블이 지정된 데이터 세트에서 뛰어난 실존/가짜 성능을 달성하며, 심각한 클래스 오염에도 안정적입니다. 또한, 일시적인 광도 곡선 클래스를 높은 정확도로 복원하지만, 단일 소스 식별은 광도 곡선에서 파생된 레이블의 모호성으로 인해 제한됩니다. 제안하는 하이브리드 불확실성 정량화 방법은 더 비용이 많이 드는 앙상블 기반 모델과 비교하여 경쟁력 있는 보정 성능을 제공합니다. 잠재 공간 분석 결과, 불확실성이 의사 결정 경계와 일치하며, 가짜 데이터 집단 내의 하위 클래스를 드러냅니다. 본 연구 결과는 주입 기반의 약한 감독 학습이 인간 레이블이 없는 훈련 데이터를 사용하지 않고도 확장 가능하고 일관된 실존/가짜 분류를 가능하게 하며, 동시에 보정된 불확실성을 제공할 수 있음을 보여줍니다. 또한, 이 방법은 향후 관측 데이터에 적용하기 위해 주입 기반의 훈련 파이프라인을 재실행하여 쉽게 이전할 수 있습니다.

Original Abstract

Time-domain surveys generate many transient candidates, making Real-Bogus classification a critical step in automated discovery pipelines. Reliable labels are costly, while community labels can be noisy and survey-dependent. We aim to develop a Real-Bogus classification framework that can be trained without human-labeled data using injected transients and bogus-dominated survey data, remains robust under strong class contamination, and provides calibrated uncertainty quantification. We combine simulated transient injections with a contaminated survey class and train a dual-network model using asymmetric co-teaching for classes with different label-noise levels. We evaluate performance on a benchmark subset and analyze the learned representation with latent-space visualization tools. For uncertainty quantification (UQ), we compare MC dropout and deep ensembles and propose a low-cost hybrid strategy that exploits the dual-network setting to improve calibration. We extend the evaluation to the light-curve domain to assess recovery of light-curve classes. The method achieves strong Real-Bogus performance on the labeled subset and remains stable under severe class contamination. It recovers transient light-curve classes with high fidelity, while single-source identification is limited by ambiguity in light-curve-derived labels. Our hybrid UQ approach achieves competitive calibration relative to more expensive ensemble baselines. Latent-space analyses indicate that uncertainty aligns with the decision boundary and reveal subclasses within the bogus population. Our results show that injection-driven, weakly supervised training can enable scalable and consistent Real-Bogus classification without human-labeled training data while providing calibrated uncertainties. The method is suited for transfer to forthcoming surveys by re-running the injection-based training pipeline.

0 Citations
0 Influential
7 Altmetric
35.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!