GPIC: 시각적 생성 모델링을 위한 거대하고 개방적인 이미지 데이터셋
GPIC: A Giant Permissive Image Corpus for Visual Generation
확장 가능한 시각적 생성 모델링 방법을 연구하기 위해서는 크고, 접근 가능하며, 안정적인 데이터셋이 필요합니다. 본 논문에서는 약 28조 픽셀로 구성된 거대한 개방형 이미지 데이터셋인 GPIC를 소개합니다. GPIC는 최첨단 컴퓨터 비전-언어 모델에 의해 설명이 추가된 다양한 인터넷 이미지를 포함하며, 학습용 1억 개의 예제, 검증용 20만 개의 예제, 그리고 테스트용 100만 개의 예제를 제공합니다. 또한, GPIC의 모든 이미지는 연구 및 상업적 용도로 사용 가능한 개방형 라이선스를 가지고 있습니다. GPIC는 안전성 필터링을 거쳤으며, 중복 데이터가 제거되었고, Hugging Face에 중앙 집중적으로 호스팅되어 있습니다. 우리는 GPIC를 활용한 생성 모델링을 위한 벤치마킹 프로토콜을 제공합니다. 마지막으로, GPIC에서의 픽셀 공간 플로우 매칭에 대한 기준 성능을 제시합니다. 저희의 데이터셋, 벤치마크, 그리고 모델은 https://huggingface.co/datasets/stanford-vision-lab/gpic 에서 이용할 수 있습니다. 평가 도구 및 코드는 https://gpic.stanford.edu 에서 제공됩니다.
Studying scalable methods for visual generative modeling requires large, accessible, and stable datasets. We introduce GPIC, a Giant Permissive Image Corpus of approximately 28 trillion pixels. GPIC comprises diverse internet images captioned by a state-of-the-art vision-language model, including 100M training, 200K validation, and 1M test examples. Moreover, all GPIC images are permissively licensed for both research and commercial use. GPIC is safety-filtered, deduplicated, and centrally hosted on Hugging Face. We provide a benchmarking protocol for generative modeling on GPIC. Finally, we provide a reference baseline for pixel-space flow matching on GPIC. Our dataset, benchmark, and models are available at https://huggingface.co/datasets/stanford-vision-lab/gpic. Evaluation toolkit and code are available at https://gpic.stanford.edu
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.