2608.06150v1 Aug 06, 2026 cs.AI

CogVis: 개방형 어휘 변화 감지 모델은 매 쿼리마다 장면을 새롭게 인식해야 하는가?

CogVis: Must Open-Vocabulary Change Detection Perceive the Scene Anew for Every Query?

Zijie Wang
Zijie Wang
Citations: 0
h-index: 0
Chengjie Zhong
Chengjie Zhong
Citations: 0
h-index: 0
Wei He
Wei He
Citations: 1,047
h-index: 17

지표면 모니터링에는 임의의 의미 범주를 식별할 수 있는 변화 감지 모델이 필요합니다. 개방형 어휘 변화 감지(OVCD)는 이러한 요구 사항을 해결합니다. 그러나 기존 방법은 종종 시간적 인식, 의미 구별 및 영역 검증을 얽히게 만들어 불안정한 결과와 중복된 계산을 초래합니다. 인간의 시각적 변화 인식을 모방하여, 우리는 인지적 기억 기반 프레임워크인 CogVis를 제안합니다. CogVis는 OVCD를 인식-기억-검증 패러다임으로 재구성하며, 먼저 장면 변화 심층 신경망(SCP)을 사용하여 고정된 양방향 특징에서 재사용 가능한, 범주에 독립적인 변화 우선 정보를 추출하여 시간적 증거와 의미 범주 결정의 분리를 가능하게 합니다. 이후, 의미 기억 조정기(SMC)는 이미지-쿼리별 의사 결정 임계값을 동적으로 추정하여 범주에 따른 점수 변동을 보정합니다. 마지막으로, 적응형 영역 필터(ARF)는 학습된 의미, 시간 및 구조적 신뢰성을 사용하여 연결된 후보들을 필터링합니다. 의미 변화 감지, 이진 변화 지역화 및 건물 손상 평가를 포함한 7개의 벤치마크에서 수행한 실험 결과, CogVis는 모든 평가 데이터 세트에서 최고 성능을 달성했습니다. CogVis는 장면 수준의 변화 인식을 공유함으로써, 쿼리마다 범주에 독립적인 시간적 인식 과정을 반복하는 것을 방지하고 추론 처리량을 28.50% 향상시킵니다.

Original Abstract

Earth-surface monitoring requires change detection models capable of recognizing arbitrary semantic categories. Open-Vocabulary Change Detection (OVCD) addresses this need. However, existing methods often entangle temporal perception, semantic discrimination, and region verification, causing unstable results and redundant computation. Inspired by human visual change perception, we propose CogVis, a cognitive memory-guided framework that reformulates OVCD as a perception-memory-verification paradigm. CogVis first employs a Scene Change Perceptron (SCP) to extract a reusable, category-agnostic change prior from frozen bi-temporal features, thereby decoupling temporal evidence from semantic category decisions. A Semantic Memory Calibrator (SMC) then compensates for category-dependent score shifts by dynamically estimating an image-query-specific decision threshold. Finally, an Adaptive Region Filter (ARF) filters connected candidates using learned semantic, temporal, and structural reliability. Experiments on seven benchmarks spanning semantic change detection, binary change localization, and building-damage assessment show that CogVis achieves state-of-the-art performance across all evaluated datasets. By sharing scene-level change perception, CogVis further avoids repeating category-agnostic temporal perception across queries and improves inference throughput by 28.50%.

0 Citations
0 Influential
8.5 Altmetric
42.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!