2607.24241v1 Jul 27, 2026 cs.CV

FilmBench: 영화 수준의 영상 생성 평가 지표

FilmBench: A Film-Grade Benchmark for Cinematic Video Generation

Fei Ding
Fei Ding
Citations: 374
h-index: 9
Wei Qiao
Wei Qiao
Citations: 53
h-index: 3
Lin Qu
Lin Qu
Citations: 164
h-index: 7
Bing Zhao
Bing Zhao
Citations: 10
h-index: 2
Guang Hu
Guang Hu
Citations: 23
h-index: 2
Hongbo Qi
Hongbo Qi
Citations: 0
h-index: 0
Xiaotong Lv
Xiaotong Lv
Citations: 0
h-index: 0
Peng Han
Peng Han
Citations: 774
h-index: 13
Fanshu Ding
Fanshu Ding
Citations: 0
h-index: 0
Mi Tang
Mi Tang
Citations: 0
h-index: 0
Hengxia Qiang
Hengxia Qiang
Citations: 0
h-index: 0
Jinyang Zhen
Jinyang Zhen
Citations: 1
h-index: 1
Jing Li
Jing Li
Citations: 0
h-index: 0
Shengyi Wang
Shengyi Wang
Citations: 0
h-index: 0
Niantong Li
Niantong Li
Citations: 12
h-index: 2
Jinlin Wang
Jinlin Wang
Citations: 803
h-index: 14
Zimeng Li
Zimeng Li
Citations: 0
h-index: 0
Han Wu
Han Wu
Citations: 99
h-index: 5
Jingjing Chen
Jingjing Chen
Citations: 0
h-index: 0
Chongxiao Wang
Chongxiao Wang
Citations: 0
h-index: 0
Yanhao Wu
Yanhao Wu
Citations: 70
h-index: 5
Cheng-Chieh Huang
Cheng-Chieh Huang
Citations: 0
h-index: 0
Xiao-quan Zhu
Xiao-quan Zhu
Citations: 0
h-index: 0
Jie Tian
Jie Tian
Citations: 13
h-index: 2
Hua Li
Hua Li
Citations: 39
h-index: 3
Jingjing Fan
Jingjing Fan
Citations: 0
h-index: 0
Zhongzhong Li
Zhongzhong Li
Citations: 0
h-index: 0
Weibin Chen
Weibin Chen
Citations: 2
h-index: 1
Huayun Wei
Huayun Wei
Citations: 0
h-index: 0
Yushu Wang
Yushu Wang
Citations: 0
h-index: 0

영상 생성 기술 발전은 AI가 생성한 영상과 전문적으로 제작된 영상 간의 시각적 격차를 줄여나가고 있지만, 대부분의 평가 지표는 여전히 웹 자료나 LLM 템플릿에서 가져온 프롬프트를 사용하고, 훈련되지 않은 범용 멀티모달 모델로 평가합니다. 더욱 근본적인 문제는 이러한 평가 기준이 전체적인 시각 품질, 대략적인 텍스트 일치성 및 시간적 부드러움과 같은 기본적인 측면에 머물러 있다는 점입니다. 이는 실제 영화 제작 및 평가에 사용되는 전문적인 영화 언어 기준과는 거리가 멀기 때문에, 영상의 기본적 타당성을 평가하는 데 그치는 것입니다. 본 연구에서는 영화 아카데미 전통에 기반한 전문적인 영화 언어를 토대로 베이징 필름 아카데미 교수진과 후징 디지털 미디어 & 엔터테인먼트 그룹 영화 스튜디오와 공동 개발한 텍스트-투-비디오(T2V) 및 레퍼런스-투-비디오(R2V) 평가 지표인 FilmBench를 소개합니다. FilmBench는 다음과 같은 세 가지 특징을 가지고 있습니다. 첫째, 프롬프트는 수상작으로 선정된 영화의 클립에서 역추적하여 생성되었으며, 전문 감독이 선별했습니다. 따라서 모든 프롬프트는 검증된 실사 레퍼런스에 연결되어 있으며, 실제 촬영 목록을 따르며, 대부분 멀티샷으로 구성되어 있습니다 (총 1,169개의 프롬프트 중 1,056개가 멀티샷). 둘째, 평가는 세 가지 수준의 영화 언어 분류 체계를 기반으로 3가지 축, 12가지 요소 및 35개(T2V) + 3개(R2V 전용) 하위 지표로 구성됩니다. 셋째, 자체 개발한 전문가 수준의 자동 평가 에이전트를 개발하고, 핵심적인 영화 언어 연산 모듈(FilmOps)을 공개합니다. FilmBench를 사용하여 선도적인 영상 생성 모델(T2V: 9개 모델, R2V: 7개 모델)을 평가한 결과, 평가 시스템은 모델 수준에서 인간의 순위와 높은 상관관계를 보였습니다 (Spearman 상관 계수: T2V = 0.95, R2V = 0.96). FilmBench의 점수는 기존 웹 기반 평가 지표보다 낮게 나타났으며, 특히 역동적인 미적 요소 및 싱글샷에서 멀티샷으로 전환할 때 성능 저하가 두드러지게 나타났습니다 (특히 성능이 낮은 모델에서 더욱 심각했습니다).

Original Abstract

Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score them with untrained, generic multimodal models. More fundamentally, their evaluation taxonomies remain rudimentary (overall visual quality, coarse text alignment and temporal smoothness) rather than the professional Cinematic Language criteria by which films are actually made and judged, so they assess basic video plausibility rather than film-grade craft. We introduce FilmBench, a text-to-video (T2V) and reference-to-video (R2V) benchmark grounded in the professional Cinematic Language of the film- academy tradition and co-developed with directors and faculty from the Beijing Film Academy and the Hujing Digital Media & Entertainment Group film studio. It rests on three choices. First, prompts are reverse-engineered from clips of award-winning films spanning 20 cinematic genres and chosen by professional directors, so every prompt is anchored to a verified live-action reference; the prompts follow real shot lists, and most script multiple shots (1,056 of the 1,169 prompts are multi-shot), unlike prior single-clip benchmarks. Second, evaluation follows a three-level Cinematic taxonomy of 3 axes, 12 components and 35 (T2V) +3 (R2V-only) sub-metrics. Third, we develop an in-house expert-grade automatic evaluation agent and open-source its core suite of Cinematic Language operators (FilmOps). Benchmarking leading video generation models (9 for T2V, 7 for R2V), the evaluator reproduces the human model ranking at model-level Spearman \r{ho} = 0.95 (T2V) and 0.96 (R2V). Scores fall well below prior web-style benchmarks, with two consistent gaps in dynamic aesthetics and a marked single- to multi-shot performance drop that widens for weaker models.

0 Citations
0 Influential
7 Altmetric
35.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!