2607.29370v1 Jul 31, 2026 cs.CV

VFAD: 변분적 의미 프롬프팅과 주파수 적응 표현 학습의 결합을 통한 제로샷 이상 탐지

VFAD: Variational Semantic Prompting Meets Frequency-Adaptive Representation Learning for Zero-Shot Anomaly Detection

Wei Wang
Wei Wang
Citations: 80
h-index: 4
Kaige Li
Kaige Li
Citations: 94
h-index: 4
Mingbo Yang
Mingbo Yang
Citations: 0
h-index: 0
Wenqiang Wang
Wenqiang Wang
Citations: 15
h-index: 2
Li Shen
Li Shen
Citations: 27
h-index: 2
Fangjun Huang
Fangjun Huang
Citations: 16
h-index: 2
Peng Chen
Peng Chen
Citations: 8
h-index: 2
Chao Huang
Chao Huang
Citations: 13
h-index: 2

제로샷 이상 탐지(ZSAD)는 특정 대상에 대한 훈련 데이터 없이, 새로운 범주에서 발생하는 이상을 탐지하고 위치시키는 것을 목표로 합니다. 최근 CLIP 기반 방법들은 시각-언어 정렬을 통해 유망한 일반화 성능을 보여주었지만, 다양한 이상 의미를 포착하고 미묘한 지역적 변화를 감지하는 데는 한계가 있습니다. 이러한 제한 사항을 해결하기 위해, 본 연구에서는 변분적 의미 프롬프팅과 주파수 적응 표현 학습을 결합한 통합 프레임워크인 VFAD를 제안합니다. 구체적으로, 우리는 밀집된 패치 토큰에서 이상 관련 지역적 의미를 능동적으로 추출하고, 변분 정보 병목 현상을 통해 이를 규제하는 Variational Semantic Prompt Extractor (VSPE)를 도입하여 미세한 시각적 단서를 통합하고 더욱 정확한 교차 모달 정렬을 가능하게 합니다. 또한, 웨이블릿 기반 주파수 분해와 주파수별 전문가 집계를 활용하여 이상 식별 능력이 뛰어난 시각적 표현을 향상시키는 Frequency-Adaptive Representation Aggregation (FARA) 모듈을 개발했습니다. VFAD는 의미 지침과 시각적 표현 학습을 동시에 강화함으로써, 이상 식별 능력과 미세한 지역 정보 위치 정확도를 모두 향상시킵니다. 13개의 산업 및 의료 벤치마크에 대한 광범위한 실험 결과, VFAD는 다양한 이상 상황에서 기존의 최첨단 ZSAD 방법보다 일관되게 우수한 성능을 보였습니다. 코드는 출판 시 공개될 예정입니다.

Original Abstract

Zero-shot anomaly detection (ZSAD) aims to detect and localize anomalies in unseen categories without access to target-specific training data. Although recent CLIP-based methods have demonstrated promising generalization through vision-language alignment, they remain limited in capturing diverse anomaly semantics and subtle local variations. To address these limitations, we propose VFAD, a unified framework that combines variational semantic prompting with frequency-adaptive representation learning. Specifically, we introduce a Variational Semantic Prompt Extractor (VSPE), which adaptively aggregates anomaly-relevant local semantics from dense patch tokens and regularizes them through a variational information bottleneck, thereby incorporating fine-grained visual cues and enabling more precise cross-modal alignment. Furthermore, we develop a Frequency-Adaptive Representation Aggregation (FARA) module that leverages wavelet-based frequency decomposition and frequency-specific expert aggregation to enhance anomaly-discriminative visual representations. By jointly strengthening semantic guidance and visual representation learning, VFAD improves both anomaly discrimination and fine-grained localization. Extensive experiments on 13 industrial and medical benchmarks demonstrate that VFAD consistently outperforms existing state-of-the-art ZSAD methods across diverse anomaly scenarios. The code will be publicly available upon publication.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!