FilmWorld: 동적 영화 세계 모델링을 통한 능동적인 소설-영화 생성
FilmWorld: Agentic Novel-to-Film Generation through Dynamic Cinematic World Modeling
소설을 영화로 변환하는 것은 생성 인공 지능 분야에서 매우 어려운 과제이며, 추상적인 문학적 텍스트를 장편의 다중 장면 시각적 내러티브로 전환해야 합니다. 현재의 비디오 생성 모델은 제한된 시간 및 공간적 맥락 내의 짧고 단일 장면 클립을 만드는 데 뛰어나지만, 소설-영화 생성은 훨씬 더 복잡한 영역에 속하며, 다양한 장면에서 동적으로 변화하는 개체의 상태를 포함하는 장기적인 콘텐츠가 필요합니다. 이러한 문제를 해결하기 위해, 우리는 소설-영화 생성을 동적 영화 세계 모델링으로 정의하고, 이를 두 단계로 구성했습니다. 첫째, 추상적이고 상세하게 설명되지 않은 문학적 내러티브를 구체적이고 상태 기반이며 지속적인 세계 개체로 변환하는 '구축(construction)' 단계입니다. 둘째, 플롯 진행에 따라 이러한 개체가 동적으로 업데이트되어 장면 간의 인과 관계 일관성을 유지하도록 하는 '진화(evolution)' 단계입니다. 우리는 FilmWorld라는 엔드투엔드 능동 시스템을 제안합니다. 이 시스템에서는 두 그룹의 전문 에이전트가 협력하여 위에서 언급한 단계를 구현합니다. 구축 단계 에이전트는 내러티브 구조 기반 번역, 시각적 고정(visual anchoring)을 통한 세계 개체 상태 모델링, 그리고 상태 기반 촬영 계획을 수행하며, 점진적으로 문학적 언어를 영화 제작 청사진으로 변환합니다. 진화 단계 에이전트는 상태 기반 시각 생성, 장면 간 동적 상태 전파, 그리고 폐루프 상태 검증을 수행하여 인과 관계 일관성과 시각적 일관성을 유지합니다. 장기 콘텐츠 생성에 대한 평가 격차를 해소하기 위해, 우리는 FilmEval이라는 체계적인 평가 프레임워크를 도입했습니다. 이 프레임워크는 15개의 대표적인 소설로 구성된 난이도 등급 벤치마크와 함께, 세 가지 차원(시네마틱 표현, 영화 일관성, 소설 충실도)에 걸쳐 9가지 객관적 지표를 포함하는 자동화된 프로토콜을 결합합니다. 실험 결과, FilmWorld는 최첨단 비디오 생성 에이전트 시스템보다 지속적으로 우수한 성능을 보이며, 특히 내러티브 충실도와 장면 간 일관성 측면에서 뚜렷한 개선 효과를 나타냅니다.
Translating novels into films poses a grand challenge for generative artificial intelligence, requiring conversion of abstract literary prose into long-form, multi-scene visual narratives. While current video generation models excel at short, single-scene clips within narrow temporal and spatial contexts, novel-to-film generation operates in a more complex regime, demanding long-duration content across diverse scenes with dynamically evolving entity states. To address this, we formalize novel-to-film generation as dynamic cinematic world modeling, decomposed into two phases: construction, which grounds abstract, underspecified literary narratives into concrete, stateful, and persistent world entities; and evolution, which governs how these entities dynamically update under plot progression to maintain causal consistency across scenes. We propose FilmWorld, an end-to-end agentic system where two groups of specialized agents collaborate to instantiate these phases. Construction-side agents perform narrative structured translation, world entity state modeling with visual anchoring, and state-driven shot planning, progressively projecting literary language into a cinematic blueprint. Evolution-side agents perform state-anchored visual generation, cross-shot dynamic state propagation, and closed-loop state verification to maintain causal consistency and visual coherence. To address the evaluation gap in long-form generation, we introduce FilmEval, a systematic evaluation framework that couples a difficulty-graded benchmark of 15 representative novels with an automated protocol of nine objective metrics spanning three dimensions: cinematic presentation, film consistency, and novel fidelity. Experiments demonstrate that FilmWorld consistently outperforms state-of-the-art video generation agent systems, with particularly pronounced improvements in narrative fidelity and cross-scene consistency.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.