2607.08076v1 Jul 09, 2026 cs.CV

LDFE: 라플라시안 분리 특징 강화 블록 - RGB-IR 이미지를 이용한 이중 스트림 CNN 기반 객체 탐지

LDFE: Laplacian Decoupled Feature Enhancement Block for Dual-Stream CNN-based RGB-IR Object Detection

Baochang Zhang
Baochang Zhang
Citations: 24
h-index: 3
Haodong Zhu
Haodong Zhu
Citations: 142
h-index: 3
Linlin Yang
Linlin Yang
Citations: 64
h-index: 5
Guodong Guo
Guodong Guo
Citations: 141
h-index: 3
Wenhao Dong
Wenhao Dong
Citations: 137
h-index: 3
Xiaoyan Luo
Xiaoyan Luo
Citations: 133
h-index: 2
Xiaorong Shi
Xiaorong Shi
Citations: 1,633
h-index: 24

RGB 이미지와 IR 이미지 간의 상호 보완적인 정보는 극한 환경에서 객체 탐지 성능을 크게 향상시킬 수 있습니다. 기존 방법들은 주로 YOLO를 기반으로 한 이중 스트림 CNN 구조를 특징 추출에 사용하며, 특징 융합 설계에 초점을 맞춥니다. 본 논문에서는 이중 스트림 CNN의 다양한 단계에서 특징을 융합하기 위해 Laplacian Decoupled Feature Enhancement (LDFE) 블록을 제안합니다. LDFE는 전역-지역 분해, 노이즈 제거, 융합 및 재구성을 순차적으로 수행함으로써, 특징 융합 과정에서 모달리티와 구조의 특성을 동시에 고려하도록 설계되었습니다. 구체적으로, LDFE는 먼저 라플라시안 피라미드를 기반으로 특징을 전역 및 지역 구성 요소로 분리하고, Global State Space Enhancement (GS2E) 모듈과 Local Convolutional Correlation Enhancement (LC2E) 모듈을 사용하여 각각 노이즈 제거 및 융합을 수행합니다. GS2E는 주 모달리티와 보조 모달리티를 위한 이중 구조를 사용하며, 보조 모달리티에서 파생된 교차 모달 어텐션을 통해 주 모달리티의 노이즈를 동적으로 억제하고, State Space 모델을 사용하여 주 모달리티의 전역 특징 표현 내 장거리 의존성을 포착합니다. 양방향 상호 작용을 위해 두 모달리티는 주/보조 역할을 체계적으로 번갈아 가며 사용합니다. 또한 LC2E는 지역 특징에서 노이즈를 억제하고, 공간 및 채널 차원과 함께 삼중 컨볼루션을 사용하여 세밀한 디테일을 추출하여 융합에 활용합니다. 이러한 혁신적인 설계 덕분에, 제안하는 방법은 M3FD, DroneVehicle, LLVIP, FLIR-Aligned, KAIST 및 VEDAI 데이터셋에서 각각 mAP 측면에서 SOTA (State-of-the-Art) 방법보다 6.2%, 3.7%, 4.7%, 2.3%, 4.1% 및 2.0%의 상당한 성능 향상을 달성했습니다.

Original Abstract

The complementary information between RGB and IR images can significantly enhance object detection performance under extreme conditions. Existing methods prefer dual-stream CNN backbones built upon YOLO for feature extraction and focus on the design of feature fusion. In this paper, we introduce the Laplacian Decoupled Feature Enhancement block (LDFE) to fuse features from different stages of the dual-stream CNN backbone. By design, LDFE simultaneously considers the characteristics of modalities and structures for feature fusion by employing global-local decomposition, denoising, fusion, and reconstruction, sequentially. The LDFE first separates features into global and local components based on Laplacian Pyramid, and then performs denoising and fusion based on Global State Space Enhancement module (GS2E) and Local Convolutional Correlation Enhancement module (LC2E) separately. Specifically, the GS2E conducts a two-branch architecture for the main and auxiliary modalities. It dynamically suppresses noise in the main modality through cross-modal attention derived from the auxiliary modality, while employing a State Space Model to capture long-range dependencies within the global feature representations of the main modality. To obtain bidirectional interaction, the two modalities systematically alternate their main/auxiliary roles. Moreover, the LC2E suppresses noise in local features and leverages spatial and channel dimension along with triple convolution to extract fine-grained details for fusion. These innovative designs achieve a significant performance improvement, with mAP surpassing the SOTA methods 6.2%, 3.7%, 4.7%, 2.3%, 4.1% and 2.0% on M3FD, DroneVehicle, LLVIP, FLIR-Aligned, KAIST and VEDAI datasets,respectively.

0 Citations
0 Influential
12 Altmetric
60.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!