CROSS: 캐스케이드 증류 및 이중 제약 기반 지상화 방법을 활용한 원격 감지 참조 분할
CROSS: Cascaded Distillation and Dual-Constraint Grounding for Remote Sensing Referring Segmentation
참조 원격 감지 이미지 분할(RRSIS)은 VLMs(Vision-Language Models) 및 Segment Anything Model (SAM)의 통합을 통해 상당한 발전을 이루었습니다. 그러나 이러한 발전은 주로 강력한 사전 훈련 능력에 의존하며, 다음 두 가지 기본적인 한계는 충분히 해결되지 않았습니다: (1) 아키텍처적 약결합(Architectural Weak-Coupling): 단방향 흐름으로 인해 VLM의 부정확한 프롬프트에 의존하게 되고 SAM의 픽셀 수준 구조 정보를 활용하지 못하여 위치 추정 오류가 발생합니다. (2) 객체 중심 의미 편향(Object-Centric Semantic Bias): 모델이 지배적인 객체 의미에 지나치게 집중하는 반면, RRSIS에 중요한 공간적 추론에는 민감하지 않습니다. 이러한 문제점을 해결하기 위해, 우리는 RRSIS를 위한 긴밀하게 통합된 새로운 패러다임인 CROSS를 제안합니다. 첫째, 아키텍처 간의 격차를 해소하기 위해 언어 정보 기반 캐스케이드 증류(LGCD)를 도입하여 SAM의 기하학적 유사성을 소프트 정규화기로 VLM 중간 레이어에 주입하고, 밀집된 구조적 정보를 활용하여 위치 추정을 개선합니다. 둘째, перспектива-공간 대비 학습(PSCL)을 통해 마스크 필터링된 오해를 유발하는 요소 및 공간-언어적 반사실 데이터를 하드 네거티브로 사용하여 교차 앵커 제약을 가함으로써 의미 단축키를 명시적으로 제거하고 진정한 논리적 일관성을 확보합니다. RRSIS 벤치마크에 대한 광범위한 실험 결과, CROSS는 최첨단 성능을 달성하며 심각한 공간 설명 변화에도 정확한 위치 추정을 유지하여 RRSIS를 위한 강력하고 새로운 패러다임을 제시합니다.
Referring Remote Sensing Image Segmentation (RRSIS) has achieved significant progress through the integration of VLMs and the Segment Anything Model (SAM). However, this progress largely relies on strong pre-trained capabilities, while leaving two fundamental limitations insufficiently addressed: (1) Architectural Weak-Coupling, where the unidirectional flow forces reliance on coarse VLM prompts and wastes SAM's pixel-level structural guidance, causing localization drift; and (2) Object-Centric Semantic Bias, where models overemphasize dominant object semantics while remaining insensitive to spatial reasoning crucial for RRSIS. Motivated by these observations, we propose CROSS, a tightly integrated paradigm for RRSIS. First, we introduce Linguistic-Guided Cascaded Distillation (LGCD) to bridge the architectural gap, which distills SAM's geometric affinities as soft regularizers into VLM intermediate layers, injecting dense structural priors to refine localization. Second, Perspective-Spatial Contrastive Learning (PSCL) imposes cross-anchored constraints by mining mask-filtered deceptive distractors and spatial-linguistic counterfactuals as hard negatives, explicitly shattering semantic shortcuts to enforce genuine logical consistency. Extensive experiments on RRSIS benchmarks demonstrate that CROSS achieves state-of-the-art performance and maintains precise localization even under severe spatial description perturbations, standing as a robust new paradigm for RRSIS.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.