2608.06943v1 Aug 07, 2026 cs.CV

범용 교차 모드 재식별을 위한 이중 공간 모달 일관성 학습

Dual-Space Modality Consistency Learning for Universal Cross-Modal Re-Identification

Bo Li
Bo Li
Citations: 20
h-index: 3
Yujian Zhao
Yujian Zhao
Citations: 17
h-index: 3
Yukang Zhao
Yukang Zhao
Citations: 0
h-index: 0
Hankun Liu
Hankun Liu
Citations: 5
h-index: 1
Haoxuan Xu
Haoxuan Xu
Citations: 14
h-index: 2
Hanzi Wan
Hanzi Wan
Citations: 0
h-index: 0
Guanglin Niu
Guanglin Niu
Citations: 44
h-index: 4

교차 모드 재식별(ReID)은 서로 다른 이미지 모달 간 동일한 개체를 식별하는 것을 목표로 하며, 가시광선-적외선 인물 ReID 및 교차 모드 선박 ReID 분야에서 널리 연구되어 왔습니다. 기존 방법들은 공간 임베딩 공간에서의 모달 일관성을 학습하여 유망한 성능을 달성했지만, 종종 고주파 표현에 나타나는 높은 차별성과 모달 민감성을 동시에 갖는 주파수 영역의 모달 불일치를 간과합니다. 또한, 대부분의 접근 방식은 특정 모달 설정에 맞춰져 있어 다양한 교차 모드 시나리오에서의 적용 가능성이 제한됩니다. 이러한 문제점을 해결하기 위해, 본 연구에서는 범용 교차 모드 ReID를 위한 이중 공간 모달 일관성 학습(DSMCL) 프레임워크를 제안합니다. 구체적으로, DSMCL은 공간 특징 분포 일관성과 주파수 영역의 차별적 일관성을 동시에 모델링합니다. 공간 모달 일관성 학습(SMCL) 브랜치는 가우시안 기반 특징 정렬을 수행하고, 주파수 정보를 고려한 차별적 일관성 학습(FDCL) 전략은 개체 정보에 민감한 교차 모드 대비 학습을 통해 고주파 표현을 규제합니다. DSMCL은 각 모달의 특성과 공유되는 개체 정보를 함께 활용하여 강력한 표현을 학습하고, 다양한 이질적인 모달 환경을 수용할 수 있는 통합 프레임워크를 구축합니다. 또한, DSMCL은 기존의 교차 모드 ReID 아키텍처에 쉽게 통합될 수 있는 플러그 앤 플레이(plug-and-play) 프레임워크입니다. SYSU-MM01, RegDB, LLCM, HOSS-ReID 및 CMShipReID 데이터셋을 사용하여 수행된 17가지 평가 프로토콜의 실험 결과, DSMCL은 여러 대표적인 기본 모델보다 일관되게 성능 향상을 보였습니다.

Original Abstract

Cross-modal Re-Identification (ReID) aims to retrieve the same identity across heterogeneous imaging modalities and has been widely studied in visible-infrared person ReID and cross-modal ship ReID. Existing methods have achieved promising performance by learning modality consistency in the spatial embedding space, yet often overlook frequency-domain modality discrepancy, particularly in high-frequency representations that are both highly discriminative and modality-sensitive. In addition, most approaches are tailored to specific modality settings, limiting their applicability across diverse cross-modal scenarios. To address these challenges, we propose a Dual-Space Modality Consistency Learning (DSMCL) framework for universal cross-modal ReID. Specifically, DSMCL jointly models spatial feature distribution consistency and frequency-domain discriminative consistency. A Spatial Modality Consistency Learning (SMCL) branch performs Gaussian-based feature alignment, while a Frequency-aware Discriminative Consistency Learning (FDCL) strategy regularizes high-frequency representations through identity-aware cross-modal contrastive learning. By jointly capturing modality-specific characteristics and modality-shared identity cues, DSMCL learns robust representations and establishes a unified framework capable of accommodating diverse heterogeneous modality settings. Moreover, DSMCL is a plug-and-play framework that can be readily integrated into existing cross-modal ReID architectures. Extensive experiments on SYSU-MM01, RegDB, LLCM, HOSS-ReID, and CMShipReID across seventeen evaluation protocols show that DSMCL consistently improves multiple representative baselines.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!