2605.06535v1 May 07, 2026 cs.CV

Sparkle: 분리된 가이드 방식을 통한 생동감 있는 지침 기반 비디오 배경 교체 구현

Sparkle: Realizing Lively Instruction-Guided Video Background Replacement via Decoupled Guidance

Guoqiang Liang
Guoqiang Liang
Citations: 120
h-index: 5
Ziyun Zeng
Ziyun Zeng
Citations: 15
h-index: 2
Yiqi Lin
Yiqi Lin
Citations: 50
h-index: 3
M. Shou
M. Shou
Citations: 13,335
h-index: 49

최근 몇 년 동안 Senorita-2M과 같은 오픈 소스 노력은 비디오 편집을 자연어 지침으로 이끌었습니다. 그러나 현재 공개적으로 사용 가능한 데이터 세트는 주로 로컬 편집 또는 스타일 변환에 중점을 두며, 이는 원본 장면 구조를 대부분 유지하고 확장하기 용이합니다. 반면, 영화 제작 및 광고와 같은 창의적인 응용 분야의 핵심 과제인 배경 교체는 정확한 전경-배경 상호 작용을 유지하면서 완전히 새로운, 시간적으로 일관된 장면을 합성해야 하므로, 대규모 데이터 생성은 훨씬 더 어렵습니다. 결과적으로, 이 복잡한 과제는 고품질 훈련 데이터의 부족으로 인해 아직 충분히 연구되지 않았습니다. 이러한 격차는 Kiwi-Edit과 같은 최첨단 모델의 성능 저하에서 명확하게 드러납니다. 왜냐하면 이 과제를 포함하는 주요 오픈 소스 데이터 세트인 OpenVE-3M이 종종 정적이고 부자연스러운 배경을 생성하기 때문입니다. 본 논문에서는 이러한 품질 저하의 원인을 데이터 합성 과정에서 정밀한 배경 가이드의 부족으로 진단했습니다. 이에 따라, 엄격한 품질 필터링을 통해 전경 및 배경 가이드를 분리된 방식으로 생성하는 확장 가능한 파이프라인을 설계했습니다. 이 파이프라인을 기반으로, 우리는 ~140,000개의 비디오 쌍으로 구성된 데이터 세트인 Sparkle과 함께, 현재까지 가장 큰 배경 교체 평가 벤치마크인 Sparkle-Bench를 소개합니다. 실험 결과, 제안된 데이터 세트와 이를 기반으로 훈련된 모델은 OpenVE-Bench와 Sparkle-Bench 모두에서 기존의 모든 모델보다 훨씬 뛰어난 성능을 보여주었습니다. 제안된 데이터 세트, 벤치마크 및 모델은 https://showlab.github.io/Sparkle/ 에서 완전히 공개적으로 제공됩니다.

Original Abstract

In recent years, open-source efforts like Senorita-2M have propelled video editing toward natural language instruction. However, current publicly available datasets predominantly focus on local editing or style transfer, which largely preserve the original scene structure and are easier to scale. In contrast, Background Replacement, a task central to creative applications such as film production and advertising, requires synthesizing entirely new, temporally consistent scenes while maintaining accurate foreground-background interactions, making large-scale data generation significantly more challenging. Consequently, this complex task remains largely underexplored due to a scarcity of high-quality training data. This gap is evident in poorly performing state-of-the-art models, e.g., Kiwi-Edit, because the primary open-source dataset that contains this task, i.e., OpenVE-3M, frequently produces static, unnatural backgrounds. In this paper, we trace this quality degradation to a lack of precise background guidance during data synthesis. Accordingly, we design a scalable pipeline that generates foreground and background guidance in a decoupled manner with strict quality filtering. Building on this pipeline, we introduce Sparkle, a dataset of ~140K video pairs spanning five common background-change themes, alongside Sparkle-Bench, the largest evaluation benchmark tailored for background replacement to date. Experiments demonstrate that our dataset and the model trained on it achieve substantially better performance than all existing baselines on both OpenVE-Bench and Sparkle-Bench. Our proposed dataset, benchmark, and model are fully open-sourced at https://showlab.github.io/Sparkle/.

0 Citations
0 Influential
24.5 Altmetric
122.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!