OpenVE-3M: 지시 기반 비디오 편집을 위한 대규모 고품질 데이터셋
OpenVE-3M: A Large-Scale High-Quality Dataset for Instruction-Guided Video Editing
지시 기반 이미지 편집 데이터셋의 품질과 다양성은 지속적으로 증가하고 있지만, 여전히 지시 기반 비디오 편집을 위한 대규모 고품질 데이터셋은 부족합니다. 이러한 격차를 해소하기 위해, 우리는 오픈 소스이며 대규모이고 고품질의 지시 기반 비디오 편집 데이터셋인 OpenVE-3M을 소개합니다. 이 데이터셋은 크게 두 가지 범주로 구성됩니다: 공간적으로 연관된 편집 (전체 스타일 변경, 배경 변경, 지역 변경, 지역 제거, 지역 추가 및 자막 편집)과 공간적으로 연관되지 않은 편집 (카메라 멀티샷 편집 및 창의적 편집). 모든 편집 유형은 엄격한 품질 필터링을 거친 정교하게 설계된 데이터 파이프라인을 통해 생성되었습니다. OpenVE-3M은 규모, 편집 유형의 다양성, 지시문의 길이 및 전반적인 품질 측면에서 기존 오픈 소스 데이터셋보다 우수합니다. 또한, 이 분야에 통일된 벤치마크가 부족하다는 점을 고려하여, 다양한 편집 작업을 포괄하는 431개의 비디오-편집 쌍으로 구성된 OpenVE-Bench를 구축했습니다. 이 벤치마크는 인간 판단과 밀접하게 일치하는 세 가지 주요 지표를 포함합니다. 우리는 또한 당사 데이터셋으로 학습된 50억 개의 파라미터를 가진 모델인 OpenVE-Edit을 제시하며, 이는 OpenVE-Bench에서 새로운 최고 성능을 달성하여 모든 기존 오픈 소스 모델 (140억 파라미터 기준 모델 포함)보다 뛰어난 효율성과 효과를 보여줍니다. 프로젝트 페이지는 https://lewandofskee.github.io/projects/OpenVE 에서 확인할 수 있습니다.
The quality and diversity of instruction-based image editing datasets are continuously increasing, yet large-scale, high-quality datasets for instruction-based video editing remain scarce. To address this gap, we introduce OpenVE-3M, an open-source, large-scale, and high-quality dataset for instruction-based video editing. It comprises two primary categories: spatially-aligned edits (Global Style, Background Change, Local Change, Local Remove, Local Add, and Subtitles Edit) and non-spatially-aligned edits (Camera Multi-Shot Edit and Creative Edit). All edit types are generated via a meticulously designed data pipeline with rigorous quality filtering. OpenVE-3M surpasses existing open-source datasets in terms of scale, diversity of edit types, instruction length, and overall quality. Furthermore, to address the lack of a unified benchmark in the field, we construct OpenVE-Bench, containing 431 video-edit pairs that cover a diverse range of editing tasks with three key metrics highly aligned with human judgment. We present OpenVE-Edit, a 5B model trained on our dataset that demonstrates remarkable efficiency and effectiveness by setting a new state-of-the-art on OpenVE-Bench, outperforming all prior open-source models including a 14B baseline. Project page is at https://lewandofskee.github.io/projects/OpenVE.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.