2605.29539v1 May 28, 2026 cs.CV

GiPL: 생성적 증강 반복형 유사 레이블링을 통한 교차 도메인 소량 데이터 객체 탐지

GiPL: Generative augmented iterative Pseudo-Labeling for Cross-Domain Few-Shot Object Detection

Yixiong Zou
Yixiong Zou
Citations: 378
h-index: 10
Yaze Zhao
Yaze Zhao
Citations: 17
h-index: 2
Shuqi Luo
Shuqi Luo
Citations: 0
h-index: 0
Yikai Qin
Yikai Qin
Citations: 26
h-index: 3
Jia-chiam Liu
Jia-chiam Liu
Citations: 16
h-index: 1
Yong Jiang
Yong Jiang
Citations: 11
h-index: 1

비전-언어 기반 모델은 교차 도메인 소량 데이터 객체 탐지(CD-FSOD)에서 뛰어난 제로샷 일반화 성능을 보여주었습니다. 그러나 이러한 모델들은 미세 조정 과정에서 두 가지 중요한 문제에 직면합니다. 첫째, 희소한 단일 인스턴스 주석으로 인해 지원 집합의 활용이 부족하고, 둘째, 극히 제한적인 대상 도메인 샘플로 인해 심각한 과적합이 발생할 수 있습니다. 이러한 문제를 해결하기 위해 본 논문에서는 효율적인 2분기 학습 프레임워크인 GiPL을 제안합니다. 첫 번째 분기에서는 반복적인 유사 레이블링 자기 학습 패러다임을 설계하여, 지원 집합에 대한 제로샷 추론을 수행하고 신뢰할 수 있는 유사 주석을 생성하며, 이를 실제 레이블과 결합하여 모델을 반복적으로 최적화함으로써 지원 집합 데이터를 최대한 활용합니다. 두 번째 분기에서는 대규모 비전-언어 모델을 사용하여 도메인 정렬된 다중 객체 주석 이미지를 합성하는 생성 데이터 증강 파이프라인을 도입하여 학습 샘플을 풍부하게 하고 과적합을 억제합니다. RUOD, CARPK, CarDD와 같은 세 가지 어려운 CD-FSOD 데이터셋에서 1/5/10-shot 환경에서의 광범위한 실험 결과, GiPL은 최첨단 방법보다 일관되게 우수한 성능을 보이며 상당한 성능 향상을 달성했습니다. 코드 및 관련 자료는 다음 링크에서 확인할 수 있습니다: [https://github.com/z-yaz/CDiscover](https://github.com/z-yaz/CDiscover)

Original Abstract

Vision-language foundation models have shown promising zero-shot generalization for Cross-Domain Few-Shot Object Detection (CD-FSOD). However, they face two critical challenges in fine-tuning: insufficient support set utilization due to sparse single-instance annotations, and severe overfitting under extremely limited target-domain samples. To address these issues, this paper proposes GiPL, an efficient two-branch training framework.In the first branch, we design an iterative pseudo-label self-training paradigm, which performs zero-shot inference on the support set to generate reliable pseudo-annotations, fuses them with ground-truth labels, and iteratively optimizes the model to fully exploit support set data. In the second branch, we introduce generative data augmentation pipeline using large vision-language models, which synthesizes domain-aligned, multi-object annotated images to enrich training samples and suppress overfitting. Extensive experiments on three challenging CD-FSOD datasets (RUOD, CARPK, CarDD) under 1/5/10-shot settings demonstrate that GiPL consistently outperforms state-of-the-art methods with significant performance gains.Code is available at \href{https://github.com/z-yaz/CDiscover}{CDiscover}.

0 Citations
0 Influential
25 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!