2608.08555v1 Aug 09, 2026 cs.CV

SC-Diff: 의미론적 교정 기반 확산 모델을 이용한 가시광선-적외선 이미지 변환

SC-Diff: Semantically Calibrated Diffusion for Visible-to-Infrared Image Translation

Ruicheng Zhang
Ruicheng Zhang
Citations: 53
h-index: 4
Deyu Meng
Deyu Meng
Citations: 115
h-index: 4
Chenqiang Gao
Chenqiang Gao
Citations: 534
h-index: 12
Junyin Zhang
Junyin Zhang
Citations: 59
h-index: 4
Siyu Huang
Siyu Huang
Harvard University
Citations: 1,490
h-index: 21
Jianxiong Ye
Jianxiong Ye
Citations: 146
h-index: 6
Haowei Gong
Haowei Gong
Citations: 6
h-index: 1

가시광선-적외선 이미지 변환은 풍부한 가시광선 이미지를 활용하여 적외선 학습 데이터를 확장하는 실용적인 방법입니다. 확산 모델은 강력한 생성 성능으로 인해 이 작업에 매우 유망합니다. 그러나 기존의 확산 기반 방법은 일반적으로 의미론적 정보를 외부 조건으로만 사용하며, 디노이징 네트워크 내의 토큰 간 상호 작용을 명시적으로 제어하지 않습니다. 그 결과, 이러한 방법들은 신뢰할 수 있는 주석 재사용에 필요한 객체의 위치, 모양 및 의미론적 레이아웃을 유지하는 데 어려움을 겪습니다. 본 연구에서는 의미론적 정보를 조건부 가이드와 내부 자기 주의(self-attention) 교정에 모두 활용하는 의미론적 교정 기반 잠재 확산 프레임워크인 SC-Diff를 제안합니다. 사전 학습된 SAM3 모델과 미리 정의된 텍스트 프롬프트를 사용하여, 먼저 가시광선 이미지에서 카테고리별 의미론적 마스크를 추출합니다. 이러한 마스크는 의미론적 맵으로 병합되어 가시광선 이미지와 함께 입력 조건으로 사용됩니다. 동일한 맵은 토큰 수준의 의미론적 레이블로 변환되어 디노이징 네트워크 내의 자기 주의를 교정하는 데 사용됩니다. 이러한 레이블을 기반으로, 우리는 Semantic-Guided Self-Attention Calibration (SGSC)라는 방법을 도입하여 동일 카테고리의 쿼리-키 쌍에 적응적으로 양의 편향을 적용합니다. 쿼리별 교정 강도는 의미론적 카테고리 간 주의 분포 및 해당 쿼리에 할당된 주의를 기준으로 결정됩니다. 원래의 주의 점수는 추가적으로 이 편향을 조절하여, 더 강한 반응을 보이는 동일 카테고리의 키에 더 큰 교정을 적용합니다. 이러한 부드러운 교정은 서로 다른 카테고리 간의 간섭을 줄이면서 전역적인 문맥적 상호 작용을 유지하여 생성된 적외선 이미지의 의미론적 일관성을 향상시킵니다. 광범위한 실험 결과, SC-Diff는 시각적 품질을 개선하고 하위 작업인 적외선 객체 감지에 더 효과적인 합성 학습 데이터를 생성함을 보여줍니다.

Original Abstract

Visible-to-infrared image translation provides a practical way to expand infrared training data using abundant visible images. Diffusion models are promising for this task because of their strong generative performance. However, existing diffusion-based methods typically use semantic priors only as external conditions, without explicitly regulating token interactions within the denoising network. Consequently, they struggle to preserve object locations, shapes, and semantic layouts required for reliable annotation reuse. We propose SC-Diff, a semantically calibrated latent diffusion framework that uses semantic priors for both conditional guidance and internal self-attention calibration. A pretrained SAM3 model with predefined text prompts first extracts category-specific semantic masks from visible images. These masks are merged into a semantic map and fused with the visible image as the input condition. The same map is converted into token-level semantic labels to calibrate self-attention in the denoising network. Based on these labels, we introduce Semantic-Guided Self-Attention Calibration (SGSC), which adaptively applies positive biases to query-key pairs of the same category. The query-wise calibration strength depends on the dispersion of attention across semantic categories and the attention assigned to the query's own category. The original attention scores further modulate the bias, giving greater calibration to same-category keys with stronger responses. This soft calibration reduces cross-category interference while retaining global contextual interactions, thereby improving semantic consistency in generated infrared images. Extensive experiments show that SC-Diff improves perceptual quality and produces more effective synthetic training data for downstream infrared object detection.

0 Citations
0 Influential
10.5 Altmetric
52.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!