2607.18850v1 Jul 21, 2026 cs.CV

OPD-IAD: 언어 판단에서 산업 이상 감지까지 - 온-정책 자기 증류를 활용한 접근

OPD-IAD: From Language Judgment to Industrial Anomaly Detection via On-Policy Self-Distillation

Zebang Cheng
Zebang Cheng
Citations: 440
h-index: 9
Wenming Yang
Wenming Yang
Citations: 24
h-index: 3
Guijin Wang
Guijin Wang
Citations: 14
h-index: 2
Jing Jin
Jing Jin
Citations: 0
h-index: 0
Nan Su
Nan Su
Citations: 86
h-index: 2
Hongbo Xu
Hongbo Xu
Citations: 23
h-index: 2
Fei Ma
Fei Ma
Citations: 67
h-index: 4
Shuimu Chen
Shuimu Chen
Citations: 13
h-index: 2

최근 대규모 시각-언어 모델(LVLM)은 이미지 수준의 이상 판단 및 해석 가능한 결함 추론을 제공하여 산업 이상 감지(IAD) 분야에서 강력한 잠재력을 보여주고 있습니다. 그러나 현재 LVLM 기반 IAD 방법은 생성된 언어 판단으로부터 정확한 픽셀 단위의 이상 지도를 생성하는 데 어려움을 겪고 있습니다. 본 연구에서는 언어가 시각적 반응을 주도하기보다는 가이드 역할을 하도록 하여, 정확한 픽셀 수준의 위치 추정을 달성하고자 합니다. 구체적으로, LVLM 기반 IAD를 위한 증거 우선 dense 온-정책 자기 증류 프레임워크인 **OPD-IAD**를 제안합니다. OPD-IAD는 모델 자체의 온-정책 판단 경로에 중요한 결함 정보를 전달하여, 최종 생성된 판단이 텍스트 답변으로만 취급되는 것이 아니라 밀집적인 지도 학습을 통해 얻어지도록 합니다. 이렇게 생성된 판단은 밀집적인 이상 감지를 위한 의미적 조건으로 사용됩니다. 이 조건을 밀집적인 시각적 증거로 변환하기 위해, **언어 기반 시각적 앵커링(Language-guided Visual Anchoring)** 방법을 도입합니다. 이는 최종 판단 조건 하에서 이미지와 질문을 재인코딩하여 의미론적 앵커를 생성하고, 이를 밀집적인 시각적 특징과 비교하는 대비 heatmap 헤드를 사용하여 이상 지도를 생성합니다. 따라서 언어 판단은 간결한 의미론적 가이드 역할을 하며, 밀집적인 시각적 특징이 여전히 픽셀 수준의 평가 기준으로 사용되어, 언어 품질이 직접적으로 픽셀 수준의 반응을 결정하지 않도록 합니다. 광범위한 실험 결과, OPD-IAD는 LVLM 기반 IAD 방법 중 가장 뛰어난 전반적인 성능을 보이며, 대부분의 이미지 수준, 픽셀 수준 및 질의 응답(QA) 지표에서 우수한 결과를 나타냅니다.

Original Abstract

Large vision-language models (LVLMs) have recently shown strong potential for industrial anomaly detection (IAD) by providing image-level anomaly judgments and interpretable defect reasoning. However, current LVLM-based IAD methods still struggle to produce precise pixel-level anomaly maps from generated language judgments. We aim to achieve precise pixel-level localization while using language as guidance rather than letting it dominate the visual response. Specifically, we propose \textbf{OPD-IAD}, an evidence-privileged dense on-policy self-distillation framework for LVLM-based IAD. OPD-IAD distills privileged defect evidence onto the model's own on-policy judgment trajectory, enabling the final generated judgment to be learned under dense supervision rather than treated only as a textual answer. The resulting judgment serves as a semantic condition for dense anomaly perception. To turn this condition into dense visual evidence, we introduce \textbf{Language-guided Visual Anchoring}, which uses a judgment reforward to re-encode the image and question under the final-judgment condition into semantic anchors and contrasts them with dense visual features through a contrastive heatmap head to generate anomaly maps. The language judgment therefore provides compact semantic guidance, while dense visual features remain the basis for pixel-level scoring, allowing language to guide anomaly localization without letting language quality directly dictate the pixel-level response. Extensive experiments show that OPD-IAD achieves the best overall performance among LVLM-based IAD methods, leading on most image-level, pixel-level, and QA metrics.

0 Citations
0 Influential
4.5 Altmetric
22.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!