MuSS: 멀티샷 피사체-비디오 생성에 대한 대규모 데이터셋 및 영화적 스토리텔링 벤치마크
MuSS: A Large-Scale Dataset and Cinematic Narrative Benchmark for Multi-Shot Subject-to-Video Generation
기존의 비디오 기반 모델은 단일 장면 생성에 탁월하지만, 실제 영화 제작은 복잡한 멀티샷 시퀀스에 의존합니다. 그러나 이러한 발전은 세 가지 핵심 과제를 해결하는 데이터셋이 부족하기 때문에 제한됩니다. 이러한 과제는 진정한 내러티브 논리, 공간-시간적 텍스트-비디오 정렬 충돌, 그리고 피사체-비디오(S2V) 생성에서 흔히 발생하는 "복사 및 붙여넣기" 문제입니다. 이러한 격차를 해소하기 위해, 우리는 멀티샷 비디오 및 S2V 생성을 위한 대규모의 이중 트랙 데이터셋인 MuSS를 소개합니다. 3,000편 이상의 영화에서 수집된 MuSS는 복잡한 몽타주 전환과 피사체 중심 내러티브를 모두 지원합니다. 이 데이터셋을 구축하기 위해, 우리는 맥락적 충돌을 제거하고 전역적인 내러티브 일관성을 확보하기 전에 로컬 수준의 정확성을 보장하는 점진적인 캡셔닝 파이프라인을 개발했습니다. 더욱 중요하게는, 우리는 S2V 생성에서 흔히 발생하는 "복사 및 붙여넣기" 문제를 근본적으로 해결하기 위한 크로스샷 매칭 메커니즘을 구현했습니다. 데이터셋과 함께, 우리는 시각적-논리 기반 패러다임과 연속적인 스토리텔링 및 3D 구조적 일관성을 엄격하게 평가하기 위한 새로운 Anti-Copy-Paste Variance (ACP-Var) 지표를 특징으로 하는 Cinematic Narrative Benchmark를 제안합니다. 광범위한 실험 결과, 현재의 기본 모델은 연속적인 내러티브 논리에 어려움을 겪거나 단순한 2D 스티커 생성기로 퇴화되는 반면, MuSS를 활용한 모델은 최첨단 수준의 내러티브 효과와 크로스샷 간 피사체 일관성을 달성함을 보여줍니다.
While video foundation models excel at single-shot generation, real-world cinematic storytelling inherently relies on complex multi-shot sequencing. Further progress is constrained by the absence of datasets that address three core challenges: authentic narrative logic, spatiotemporal text-video alignment conflicts, and the "copy-paste" dilemma prevalent in Subject-to-Video (S2V) generation. To bridge this gap, we introduce MuSS, a large-scale, dual-track dataset tailored for multi-shot video and S2V generation. Sourced from over 3,000 movies, MuSS explicitly supports both complex montage transitions and subject-centric narratives. To construct this dataset, we pioneer a progressive captioning pipeline that eliminates contextual conflicts by ensuring local shot-level accuracy before enforcing global narrative coherence. Crucially, we implement a cross-shot matching mechanism to fundamentally eradicate the S2V copy-paste shortcut. Alongside the dataset, we propose the Cinematic Narrative Benchmark, featuring a visual-logic-driven paradigm and a novel Anti-Copy-Paste Variance (ACP-Var) metric to rigorously assess continuous storytelling and 3D structural consistency. Extensive experiments demonstrate that while current baselines struggle with continuous narrative logic or degenerate into trivial 2D sticker generators, our MuSS-augmented model achieves state-of-the-art narrative effectiveness and cross-shot identity preservation.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.