2606.23611v1 Jun 22, 2026 cs.CV

비전-언어 모델 학습을 위한 반복적 자기 필터링 기반 데이터 선택 방법

Data Selection Through Iterative Self-Filtering for Vision-Language Settings

Aaron C. Courville
Aaron C. Courville
Citations: 212
h-index: 8
Andrei Liviu Nicolicioiu
Andrei Liviu Nicolicioiu
Mila
Citations: 127
h-index: 5
Sarvjeet Ghotra
Sarvjeet Ghotra
Citations: 414
h-index: 2
Morgane M Moss
Morgane M Moss
Citations: 29
h-index: 2

신경망 학습에는 양질의 대규모 데이터 확보가 매우 중요합니다. 그러나 데이터 규모가 커지면 수동 검토가 비현실적이 되어, 상당한 수준의 노이즈를 포함하는 데이터셋이 생성될 수 있습니다. 이러한 문제를 해결하고 우수한 성능을 보이는 비전-언어 모델을 개발하기 위한 기존 시도들은 휴리스틱 기법, 선별된 참조 데이터셋, 그리고 사전 학습된 모델 활용 등에 초점을 맞춰왔습니다. 본 연구에서는 CLIP 모델을 사용하여 진화하는 자기 선택 데이터셋으로 학습하는 새로운 방법을 제안합니다. 이 진화하는 데이터셋은 필터링되어 높은 확률로 깨끗한 샘플과 전체 분포에서 추출된 다양한 샘플의 균형을 이루도록 구성됩니다. 제안하는 자기 필터링 방법은 모델 학습과 이후 개선된 데이터 혼합 선택 단계를 반복적으로 수행합니다. 제안하는 방식으로 필터링된 비전-언어 데이터를 사용하여 모델을 훈련하면 추가적인 데이터나 사전 학습된 모델 없이도 하위 작업의 성능을 향상시킬 수 있습니다.

Original Abstract

The availability of large amounts of clean data is paramount to training neural networks. However, at large scales, manual oversight is impractical, resulting in sizeable datasets that can be very noisy. Attempts to mitigate this obstacle to producing performant vision-language models have so far involved heuristics, curated reference datasets, and using pre-trained models. Here we propose a novel, bootstrapped method in which a CLIP model is trained on an evolving, self-selected dataset. This evolving dataset constitutes a balance of filtered, highly probable clean samples as well as diverse samples from the entire distribution. Our proposed Self-Filtering method iterates between training the model and selecting a subsequently improved data mixture. Training on vision-language datasets filtered by the proposed approach improves downstream performance without the need for additional data or pre-trained models.

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!