2608.04394v1 Aug 05, 2026 cs.CV

확산 모델 기반 데이터 생성을 재검토하여 교차 도메인 소량 샘플 객체 탐지 성능 향상: Free-Lunch Augmentation

Free-Lunch Augmentation by Revisiting Diffusion-Based Data Generation for Cross-Domain Few-Shot Object Detection

Yixiong Zou
Yixiong Zou
Citations: 378
h-index: 10
Yuhua Li
Yuhua Li
Citations: 333
h-index: 10
Ruixuan Li
Ruixuan Li
Citations: 450
h-index: 12
Zijian Zhuang
Zijian Zhuang
Citations: 56
h-index: 3

교차 도메인 소량 샘플 객체 탐지(CDFSOD)는 풍부한 데이터를 가진 상위 도메인에서 지식을 하위 전문 도메인으로 이전하는 것을 목표로 하며, 제한된 학습 데이터와 큰 도메인 간 격차로 인해 해결되지 않은 과제입니다. 이 문제를 해결하기 위해, 우리는 CDFSOD에서 간과되어 왔지만 자연스러운 접근 방식인 데이터 증강을 재검토합니다. 구체적으로, 확산 모델을 통해 데이터를 직접 합성하여 제한적인 학습 샘플을 보충합니다. 그러나 도메인 간 격차가 크기 때문에, 기존의 확산 방법은 좋은 결과를 생성하지 못하며, 심지어 원본 이미지보다 성능이 낮아지는 현상이 발생합니다. 이러한 한계를 극복하기 위해, 우리는 시각적 격차와 의미적 격차로 나누어 분석을 진행했습니다. 시각적 격차에 대해서는, 확산 모델이 전문 도메인에서 유용한 정보와 노이즈를 구별하지 못한다는 것을 발견했으며, 이를 약화된 노이즈 추가를 통해 완화할 수 있습니다. 의미적 격차에 대해서는, 배경의 의미론적 격차가 객체의 의미론적 격차보다 작다는 것을 확인했고, 배경 채우기(inpainting)를 통해 이 격차를 해소할 수 있습니다. 위 분석을 바탕으로, 우리는 일반 도메인과의 격차에 따라 다양한 전략을 동적으로 적용하는 데이터 합성 방법(Selective Inpainting with Tailored Noise, SITN)을 제안합니다. 여기에는 맞춤형 노이즈 추가를 위한 생성 모듈과 채우기 영역을 동적으로 선택하는 선택 모듈이 포함됩니다. CDFSOD 6개 데이터셋 및 교차 도메인 소량 샘플 분할(CDFSS) 4개 데이터셋에 대한 광범위한 실험 결과, 유용한 데이터를 합성하여 새로운 최고 성능을 달성할 수 있음을 확인했습니다. 저희 코드는 https://github.com/zzzzj311-droid/Free-Lunch-SITN 에서 확인할 수 있습니다.

Original Abstract

Cross-Domain Few-Shot Object Detection (CDFSOD) aims to transfer knowledge from data-rich upstream generic domains to downstream expert domains using scarce training data, where the significant domain gap and data scarcity make it an unsolved challenge. To address this problem, we revisit a natural yet underexplored approach in CDFSOD: data augmentation, by directly synthesizing data through diffusion models to supplement limited training samples. However, due to large domain gaps, we find that current diffusion methods cannot produce good results, leading to performance even lower than using the original images. To address these limitations, we divide the domain gaps into visual gaps and semantic gaps for separate analysis. For the visual gap, we find that the diffusion model cannot distinguish noise from useful information on expert domains, which can be mitigated by adding weakened noise. For the semantic gap, we find that the background semantics shows much smaller gaps between domains than foreground semantics, and we can bridge this gap by background inpainting. Based on the above analysis, we propose a method (Selective Inpainting with Tailored Noise, SITN) to dynamically take different strategies for downstream data synthesis based on their different gaps from the general domain, including a Generation Module for adding tailored noise and a Selection Module to dynamically select the inpainting regions. Extensive experiments on 6 datasets of CDFSOD and 4 datasets of cross-domain few-shot segmentation (CDFSS) validate that we can synthesize helpful data, achieving new state-of-the-art performance. Our codes is available at https://github.com/zzzzj311-droid/Free-Lunch-SITN

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!