의사 라벨 기반 생성을 통한 표 데이터 이상 탐지 성능 향상
Enhancing Tabular Anomaly Detection via Pseudo-Label-Guided Generation
표 데이터에서 이상 데이터를 식별하는 것은 데이터 신뢰성을 향상시키고 시스템 안정성을 유지하는 데 필수적입니다. 실제 이상 데이터에 대한 라벨이 부족하기 때문에, 기존 방법들은 주로 비지도 이상 탐지 모델에 의존하거나, 소수의 라벨링된 이상 데이터를 활용하여 샘플 생성 또는 대비 학습을 통해 탐지를 돕습니다. 그러나 비지도 방법은 충분한 이상 인지 능력이 부족하며, 현재의 생성 및 대비 학습 방식은 전반적인 이상을 계산하는 경향이 있어 표 데이터의 지역적인 이상 패턴을 간과하게 되어, 최적의 탐지 성능을 달성하지 못합니다. 이러한 한계를 해결하기 위해, 표 데이터 이상 탐지 성능을 향상시키기 위한 의사 라벨 기반 생성 방법인 PLAG를 제안합니다. PLAG는 의사 이상 데이터를 안내 신호로 활용하고, 샘플의 전반적인 이상 정도를 특징 수준의 이상으로 분해하여, 실제 라벨이 부족한 상황에서도 효과적으로 대처할 뿐만 아니라, 모델이 미세한 수준에서 지역적인 이상 신호를 이해할 수 있도록 하는 새로운 관점을 제공합니다. 또한, 데이터 정합성 검증 및 불확실성 추정을 통합한 두 단계의 데이터 선택 전략을 제안하여, 후보 샘플을 엄격하게 필터링함으로써 생성된 이상 데이터의 신뢰성과 다양성을 보장합니다. 궁극적으로, 필터링된 생성 데이터는 강력한 판별력 있는 안내 역할을 수행하여 모델이 정상 데이터와 이상 데이터를 더 잘 구별할 수 있도록 합니다. 광범위한 실험 결과, PLAG가 8가지 대표적인 기존 방법보다 뛰어난 성능을 달성하는 것을 보여줍니다. 또한, PLAG는 유연한 프레임워크로서 기존의 비지도 탐지 시스템과 원활하게 통합되어 F1 점수를 지속적으로 0.08에서 0.21까지 향상시킵니다.
Identifying anomalous instances in tabular data is essential for improving data reliability and maintaining system stability. Due to the scarcity of ground-truth anomaly labels, existing methods mainly rely on unsupervised anomaly detection models, or exploit a small number of labeled anomalies to facilitate detection via sample generation or contrastive learning. However, unsupervised methods lack sufficient anomaly awareness, while current generation and contrastive approaches tend to compute anomalies globally, overlooking the localized anomaly patterns of tabular features, resulting in suboptimal detection performance. To address these limitations, we propose PLAG, a pseudo-label-guided anomaly generation method designed to enhance tabular anomaly detection. Specifically, by utilizing pseudo-anomalies as guidance signals and decoupling the overall anomaly quantification of a sample into an accumulation of feature-level abnormalities, PLAG not only effectively obviates the need for scarce ground-truth labels but also provides a novel perspective for the model to comprehend localized anomalous signals at a fine-grained level. Furthermore, a two-stage data selection strategy is proposed, integrating format verification and uncertainty estimation to rigorously filter candidate samples, thereby ensuring the fidelity and diversity of the synthetic anomalies. Ultimately, these filtered synthetic anomalies serve as robust discriminative guidance, empowering the model to better separate normal and anomalous instances. Extensive experiments demonstrate that PLAG achieves state-of-the-art performance against eight representative baselines. Moreover, as a flexible framework, it integrates seamlessly with existing unsupervised detectors, consistently boosting F1-scores by 0.08 to 0.21.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.