FilmBench: 영화 수준의 영상 생성 평가 지표
FilmBench: A Film-Grade Benchmark for Cinematic Video Generation
영상 생성 기술 발전은 AI가 생성한 영상과 전문적으로 제작된 영상 간의 시각적 격차를 줄여나가고 있지만, 대부분의 평가 지표는 여전히 웹 자료나 LLM 템플릿에서 가져온 프롬프트를 사용하고, 훈련되지 않은 범용 멀티모달 모델로 평가합니다. 더욱 근본적인 문제는 이러한 평가 기준이 전체적인 시각 품질, 대략적인 텍스트 일치성 및 시간적 부드러움과 같은 기본적인 측면에 머물러 있다는 점입니다. 이는 실제 영화 제작 및 평가에 사용되는 전문적인 영화 언어 기준과는 거리가 멀기 때문에, 영상의 기본적 타당성을 평가하는 데 그치는 것입니다. 본 연구에서는 영화 아카데미 전통에 기반한 전문적인 영화 언어를 토대로 베이징 필름 아카데미 교수진과 후징 디지털 미디어 & 엔터테인먼트 그룹 영화 스튜디오와 공동 개발한 텍스트-투-비디오(T2V) 및 레퍼런스-투-비디오(R2V) 평가 지표인 FilmBench를 소개합니다. FilmBench는 다음과 같은 세 가지 특징을 가지고 있습니다. 첫째, 프롬프트는 수상작으로 선정된 영화의 클립에서 역추적하여 생성되었으며, 전문 감독이 선별했습니다. 따라서 모든 프롬프트는 검증된 실사 레퍼런스에 연결되어 있으며, 실제 촬영 목록을 따르며, 대부분 멀티샷으로 구성되어 있습니다 (총 1,169개의 프롬프트 중 1,056개가 멀티샷). 둘째, 평가는 세 가지 수준의 영화 언어 분류 체계를 기반으로 3가지 축, 12가지 요소 및 35개(T2V) + 3개(R2V 전용) 하위 지표로 구성됩니다. 셋째, 자체 개발한 전문가 수준의 자동 평가 에이전트를 개발하고, 핵심적인 영화 언어 연산 모듈(FilmOps)을 공개합니다. FilmBench를 사용하여 선도적인 영상 생성 모델(T2V: 9개 모델, R2V: 7개 모델)을 평가한 결과, 평가 시스템은 모델 수준에서 인간의 순위와 높은 상관관계를 보였습니다 (Spearman 상관 계수: T2V = 0.95, R2V = 0.96). FilmBench의 점수는 기존 웹 기반 평가 지표보다 낮게 나타났으며, 특히 역동적인 미적 요소 및 싱글샷에서 멀티샷으로 전환할 때 성능 저하가 두드러지게 나타났습니다 (특히 성능이 낮은 모델에서 더욱 심각했습니다).
Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score them with untrained, generic multimodal models. More fundamentally, their evaluation taxonomies remain rudimentary (overall visual quality, coarse text alignment and temporal smoothness) rather than the professional Cinematic Language criteria by which films are actually made and judged, so they assess basic video plausibility rather than film-grade craft. We introduce FilmBench, a text-to-video (T2V) and reference-to-video (R2V) benchmark grounded in the professional Cinematic Language of the film- academy tradition and co-developed with directors and faculty from the Beijing Film Academy and the Hujing Digital Media & Entertainment Group film studio. It rests on three choices. First, prompts are reverse-engineered from clips of award-winning films spanning 20 cinematic genres and chosen by professional directors, so every prompt is anchored to a verified live-action reference; the prompts follow real shot lists, and most script multiple shots (1,056 of the 1,169 prompts are multi-shot), unlike prior single-clip benchmarks. Second, evaluation follows a three-level Cinematic taxonomy of 3 axes, 12 components and 35 (T2V) +3 (R2V-only) sub-metrics. Third, we develop an in-house expert-grade automatic evaluation agent and open-source its core suite of Cinematic Language operators (FilmOps). Benchmarking leading video generation models (9 for T2V, 7 for R2V), the evaluator reproduces the human model ranking at model-level Spearman \r{ho} = 0.95 (T2V) and 0.96 (R2V). Scores fall well below prior web-style benchmarks, with two consistent gaps in dynamic aesthetics and a marked single- to multi-shot performance drop that widens for weaker models.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.