2603.19608v1 Mar 20, 2026 cs.CV

FB-CLIP: 전경-배경 분리를 통한 정밀한 제로샷 이상 탐지

FB-CLIP: Fine-Grained Zero-Shot Anomaly Detection with Foreground-Background Disentanglement

Ming Hu
Ming Hu
Citations: 4
h-index: 1
Yongsheng Huo
Yongsheng Huo
Citations: 2
h-index: 1
Mingyu Dou
Mingyu Dou
Citations: 9
h-index: 2
Jianfu Yin
Jianfu Yin
Citations: 13
h-index: 3
Peng Zhao
Peng Zhao
Citations: 35
h-index: 4
Yao Wang
Yao Wang
Citations: 5
h-index: 1
Cong Hu
Cong Hu
Citations: 4
h-index: 1
Bingliang Hu
Bingliang Hu
Citations: 242
h-index: 8
Quan Wang
Quan Wang
Citations: 10
h-index: 2

정밀한 이상 탐지는 산업 및 의료 분야에서 매우 중요하지만, 레이블이 있는 이상 데이터가 부족한 경우가 많아 제로샷 탐지가 어렵습니다. CLIP과 같은 비전-언어 모델은 유망한 해결책을 제시하지만, 전경-배경 특징의 얽힘과 거친 텍스트 의미론으로 인해 어려움을 겪습니다. 우리는 다중 전략 텍스트 표현과 전경-배경 분리를 통해 이상 위치 파악을 향상시키는 프레임워크인 FB-CLIP을 제안합니다. 텍스트 모달리티에서는 End-of-Text 특징, 전역 풀링 표현, 그리고 어텐션 가중 토큰 특징을 결합하여 풍부한 의미적 정보를 제공합니다. 시각 모달리티에서는 동일성, 의미, 그리고 공간 차원을 따라 소프트 분리를 수행하고 배경 억제를 통해 간섭을 줄이고 구별력을 향상시킵니다. 의미 일관성 정규화(SCR)는 이미지 특징을 정상 및 비정상 텍스트 프로토타입과 정렬하여 불확실한 매칭을 억제하고 의미적 간극을 확대합니다. 실험 결과, FB-CLIP은 복잡한 배경에서 이상을 효과적으로 구별하여 제로샷 환경에서 정확한 정밀한 이상 탐지 및 위치 파악을 달성하는 것을 보여줍니다.

Original Abstract

Fine-grained anomaly detection is crucial in industrial and medical applications, but labeled anomalies are often scarce, making zero-shot detection challenging. While vision-language models like CLIP offer promising solutions, they struggle with foreground-background feature entanglement and coarse textual semantics. We propose FB-CLIP, a framework that enhances anomaly localization via multi-strategy textual representations and foreground-background separation. In the textual modality, it combines End-of-Text features, global-pooled representations, and attention-weighted token features for richer semantic cues. In the visual modality, multi-view soft separation along identity, semantic, and spatial dimensions, together with background suppression, reduces interference and improves discriminability. Semantic Consistency Regularization (SCR) aligns image features with normal and abnormal textual prototypes, suppressing uncertain matches and enlarging semantic gaps. Experiments show that FB-CLIP effectively distinguishes anomalies from complex backgrounds, achieving accurate fine-grained anomaly detection and localization under zero-shot settings.

1 Citations
0 Influential
4 Altmetric
21.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!