VideoCoCo: 물리적으로 일관된 비디오 생성을 위한 에이전트 기반 이중 엔진 시스템에서의 코드-as-CoT
VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System
텍스트-투-비디오 모델은 놀라운 시각적 품질을 달성했지만, 장면의 시간적 변화가 압축된 텍스트 프롬프트로부터 암묵적으로 추론되어야 하기 때문에 물리적으로 일관된 움직임을 생성하는 데 어려움을 겪습니다. 기존의 체인-오브-생상(CoT) 접근 방식은 중간 계획 또는 시각적 상태를 도입하지만, 이러한 표현은 일반적으로 실행 불가능하거나 시간적으로 희소하여 완전한 공간-시간 프로세스를 구현하고 제어하는 능력을 제한합니다. 이 한계를 해결하기 위해, 우리는 실행 가능한 Blender 코드를 프로세스 수준의 CoT로 사용하는 에이전트 기반 이중 엔진 프레임워크인 VideoCoCo를 소개합니다. 주어진 텍스트 프롬프트에 대해, 코딩 에이전트는 장면과 시간적 변화를 명시적으로 지정하는 Blender 프로그램을 생성합니다. 실행 가능한 시뮬레이션 엔진은 이 프로그램을 실행하여 결정적인 공간-시간 초안을 생성하며, 이는 생성 비디오 엔진에 의해 초안 조건 편집을 통해 사실적인 비디오로 변환됩니다. 이러한 분리는 프로세스 수준의 추론과 고품질 시각적 구현을 분리합니다. 시뮬레이션된 초안에 적응하도록 비디오 편집기를 조정하기 위해, 우리는 초안-지침-대상 3중 구조로 구성된 데이터셋인 VideoCoCo-3K를 구축했습니다. VideoCoCo는 OmniWeaving의 기본 성능을 PhyGenBench에서 0.475에서 0.558로, VBench-2.0에서 52.18에서 77.88로 향상시켜 두 벤치마크 모두에서 최고의 평균 점수를 달성했습니다. 이러한 결과는 실행 가능한 코드가 물리적으로 일관된 비디오 생성을 위한 효과적이고 제어 가능하며 검사 가능한 중간 표현을 제공한다는 것을 보여줍니다.
Text-to-video models have achieved remarkable visual quality, yet they still struggle to generate physically consistent dynamics because the temporal evolution of a scene must be inferred implicitly from a highly compressed text prompt. Existing chain-of-thought approaches introduce intermediate plans or visual states, but these representations are typically non-executable or temporally sparse, limiting their ability to instantiate and control the complete spatiotemporal process. To address this limitation, we introduce VideoCoCo, an agentic dual-engine framework in which executable Blender code serves as a process-level chain of thought. Given a text prompt, a coding agent synthesizes a Blender program that explicitly specifies the scene and its temporal evolution. The executable simulation engine runs the program to produce a deterministic spatiotemporal draft, which is subsequently transformed into a photorealistic video by a generative video engine through draft-conditioned editing. This decomposition separates process-level reasoning from high-fidelity visual realization. To adapt the video editor to simulated drafts, we construct VideoCoCo-3K, a curated dataset of draft-instruction-target triplets. VideoCoCo improves the OmniWeaving baseline from 0.475 to 0.558 on PhyGenBench and from 52.18 to 77.88 on VBench-2.0, achieving the best average score on both benchmarks. These results demonstrate that executable code provides an effective, controllable, and inspectable intermediate representation for physically consistent video generation.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.