LogiShot: 논리적으로 일관된 장면 간 비디오 생성
LogiShot: Logically Coherent Cross-Shot Video Generation
논리적으로 연결된 장면으로 구성된 비디오를 생성하는 것은 콘텐츠 제작에 필수적입니다. 현재 대부분의 장면 간 비디오 생성 워크플로우, 예를 들어 단편 드라마 제작은 여전히 분리된 텍스트 스크립트 또는 명시적인 참조 이미지를 사용하여 생성될 내용을 지정합니다. 결과적으로 사용자 지침이 불충분하거나 모호할 경우, 생성된 클립은 자체적으로는 시각적으로 타당해 보일 수 있지만 전체 스토리와 일치하지 않아 내용의 단절을 초래할 수 있습니다. 우리는 장면 간 비디오 생성에서 논리적 일관성을 달성하려면 장면 간 논리적 연결을 확립하고 시각적 일관성을 유지해야 한다고 주장합니다. 이를 위해, 우리는 LogiShot을 제안하며, 이 모델은 다음과 같은 두 가지 상호 보완적인 방식으로 정보를 활용합니다. 1) LogiShot은 컨텍스트 비디오 및 기타 조건부 신호를 동시에 인코딩하여 장면 간 생성을 위한 시각-의미 정보를 제공하는 풍부한 다중 모드 단서를 생성합니다. 2) 모델은 생성 프로세스 전반에 걸쳐 컨텍스트 비디오의 시각적 기억을 유지하여 장면 전체의 시각적 일관성을 보장합니다. 또한, 우리는 11만 개의 샘플로 구성된 데이터셋과 장면 간 논리적 일관성을 평가하기 위한 특별한 벤치마크를 구축했습니다. 실험 결과, LogiShot은 여러 장면에서 논리적 일관성 측면에서 기존 모델보다 우수한 성능을 보였습니다. 모델 및 데이터는 공개적으로 제공될 예정입니다.
Generating cross-shot videos that are logically connected is essential for content creation. Currently, most cross-shot video-generation workflows, such as short-drama production, still rely on isolated textual scripts or explicit reference images to specify the generated content. Consequently, when user instructions are underspecified or ambiguous, a generated clip may appear visually plausible on its own but fail to align with the overall narrative, leading to disjointed content. We argue that achieving cross-shot logical coherence in video generation requires establishing logical connections across shots and maintaining visual consistency. To this end, we propose LogiShot, which incorporates information through two complementary paths: 1) LogiShot jointly encodes the context video and other conditioning signals, yielding dense multimodal cues that provide visual-semantic evidence for cross-shot generation; 2) the model maintains a visual memory of the context video throughout generation to preserve visual consistency across shots. Additionally, we construct a dataset with 110K samples and a dedicated benchmark for evaluating cross-shot logical coherence. Experiments demonstrate that LogiShot consistently outperforms existing baselines in terms of logical coherence across multiple shots. Model and data will be made publicly available.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.