2606.17020v1 Jun 15, 2026 cs.CV

FusionRS: 다중 모드 비전-언어 기초 모델을 위한 대규모 RGB-적외선 원격 감지 데이터셋

FusionRS: A Large-Scale RGB-Infrared Remote Sensing Dataset for Dual-Modal Vision-Language Foundation Models

Jiujiang Guo
Jiujiang Guo
Citations: 73
h-index: 5
Yiwei Wei
Yiwei Wei
Citations: 360
h-index: 12
Chengyin Hu
Chengyin Hu
Citations: 6
h-index: 1
Yuxiang Dong
Yuxiang Dong
Citations: 50
h-index: 2
Xu Sun
Xu Sun
Citations: 99
h-index: 3
Jiajun Han
Jiajun Han
Citations: 0
h-index: 0
Qike Zhang
Qike Zhang
Citations: 22
h-index: 3
Benqi Zhang
Benqi Zhang
Citations: 1
h-index: 1
Fengyu Zhang
Fengyu Zhang
Citations: 2
h-index: 1
Dingyi Lu
Dingyi Lu
Citations: 0
h-index: 0
Luwei Yang
Luwei Yang
Citations: 0
h-index: 0

원격 감지 비전-언어 모델은 지구 관찰에 대한 이해를 향상시키지만, 대부분의 기존 연구는 여전히 RGB 이미지에 집중되어 있으며 적외선 데이터에서 얻을 수 있는 상호 보완적인 정보가 충분히 활용되지 못하고 있습니다. 적외선 이미지는 열 강도 구조, 객체 경계 및 조명 변화에 덜 민감한 장면 특징과 같은 고유한 정보를 제공하며, 이는 기존 RGB 관찰을 넘어 시각-언어 학습을 풍부하게 할 수 있습니다. 그러나 원격 감지 비전-언어 모델링을 위한 대규모의 RGB-적외선-텍스트 데이터셋은 아직 존재하지 않습니다. 이러한 격차를 해소하기 위해, 우리는 다중 모드 비전-언어 학습을 위한 최초의 대규모 RGB-적외선-텍스트 데이터셋인 FusionRS를 소개합니다. FusionRS는 다양한 공개 RGB 원격 감지 이미지를 적외선 스타일 이미지로 변환하여 정렬된 RGB-IR 이미지 쌍을 구성하여 구축되었습니다. 각 쌍은 기존 장면 캡션과 함께, 의미 내용을 유지하면서 적외선 특유의 시각적 특징을 명시적으로 설명하는 IR 인식 캡션을 포함합니다. FusionRS를 기반으로, 우리는 RGB-IR 공동 이해를 위한 다중 모드 비전-언어 기초 모델을 학습했습니다. 먼저 CLIP 스타일 모델을 사용하여 RGB-IR-텍스트 정렬을 수행하고, 그 다음 생성형 VLMs를 미세 조정하여 다중 모드 RGB-IR 캡셔닝을 구현합니다. 실험 결과는 FusionRS가 RGB-IR 정렬, 적외선-텍스트 검색 및 다중 모드 캡셔닝 성능을 향상시키며, 이는 RGB 전용 또는 IR 인식 학습 환경보다 우수함을 보여줍니다. 추가 분석 연구를 통해 IR 인식 캡션이 적외선-언어 정렬을 강화하는 데 중요한 역할을 한다는 것을 확인했으며, 이는 더욱 확장 가능한 RGB-적외선 원격 감지 비전-언어 표현 학습을 위한 모달리티별 텍스트 감독의 중요성을 강조합니다.

Original Abstract

Remote sensing vision-language models have advanced Earth observation understanding, but most existing work remains centered on RGB imagery, leaving the complementary information in infrared data underexplored. Infrared images provide distinctive cues, including thermal intensity structures, object boundaries, and illumination-invariant scene features, which can enrich visual-language learning beyond conventional RGB observations. However, a large-scale RGB-infrared-text dataset for remote sensing vision-language modeling is still absent. To address this gap, we introduce FusionRS, the first large-scale RGB-infrared-text dataset designed for dual-modal vision-language learning in remote sensing. FusionRS is constructed by translating diverse public RGB remote sensing images into infrared-style counterparts, forming aligned RGB-IR image pairs. Each pair is associated with conventional scene captions and IR-aware captions that explicitly describe infrared-specific visual properties while preserving semantic content. Based on FusionRS, we train dual-modal vision-language foundation models for RGB-IR joint understanding. We first train CLIP-style models for RGB-IR-text alignment, and then fine-tune generative VLMs for dual-modal RGB-IR captioning. Experiments show that FusionRS improves RGB-IR alignment, infrared-to-text retrieval, and dual-modal captioning over RGB-only and non-IR-aware training settings. Ablation studies further verify that IR-aware captions are crucial for strengthening infrared-language alignment, highlighting the importance of modality-specific textual supervision for more scalable RGB-infrared remote sensing vision-language representation learning.

1 Citations
0 Influential
6 Altmetric
31.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!