2608.09873v1 Aug 10, 2026 cs.CV

Sci-VBench: 과학 분야의 지식 및 추론 기반 비디오 생성 평가

Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Tingyu Song
Tingyu Song
University of Chineses Academy of Sciences
Citations: 32
h-index: 4
Linbo Fu
Linbo Fu
Citations: 4
h-index: 1
Zheyuan Yang
Zheyuan Yang
Citations: 20
h-index: 1

본 논문에서는 과학 분야 전반에 걸쳐 지식과 추론 능력을 요구하는 비디오 생성을 평가하기 위한 종합적인 벤치마크인 Sci-VBench를 소개합니다. 이 벤치마크는 천연과학, 의료, 인문사회과학, 공학의 네 가지 주요 학문 분야에 걸쳐 60개의 주제를 다루는 1,253개의 전문가가 직접 작성한 예시로 구성되어 있습니다. 각 예시는 모델이 과학적 추론과 지식 기반 종합을 요구하는 시간적으로 풍부한 비디오를 생성하도록 요구하며, 단순한 시각적 타당성을 넘어섭니다. 또한, 명확한 평가 기준에 따른 평가 프로토콜을 수립했습니다. 분석 결과, 본 프로토콜 하에서 비전문가 평가자와 MLLM-as-Judge 시스템 모두 전문가의 판단과 비교적 높은 일치도를 보였으며, 이는 대규모 환경에서의 재현 가능한 평가를 지원합니다. 16개의 최첨단 독점 및 오픈 소스 모델을 벤치마킹한 결과, 자동화된 시각 품질 점수는 시스템 간에 밀접하게 분포하는 반면, 프롬프트 정확성 및 과학적/인과 관계의 정확성은 상당한 차이를 보였으며, 특히 독점 모델과 오픈 소스 모델 간에 두드러진 격차가 있었습니다. 이러한 결과는 시각적 사실감 개선이 아직까지 과학적 및 인과 관계 역학을 신뢰성 있게 모델링하는 데로 이어지지 않았음을 보여줍니다.

Original Abstract

We introduce Sci-VBench, a comprehensive benchmark for evaluating knowledge- and reasoning-intensive video generation across scientific domains. It contains 1,253 expert-annotated examples spanning 60 subjects across four core disciplines: Natural Science, Healthcare, Humanities & Social Sciences, and Engineering. Each example requires models to generate temporally rich videos that demand scientific reasoning and knowledge-grounded synthesis, going beyond surface-level visual plausibility. We further establish a rubric-based evaluation protocol. Our analysis shows that, under this protocol, both non-expert human evaluators and MLLM-as-Judge systems can achieve relatively high agreement with expert judgments, supporting reproducible evaluation at scale. We benchmark 16 frontier proprietary and open-source models and find that, while automatic perceptual-quality scores cluster tightly across systems, performance on Prompt Grounding and Scientific and Causal Correctness varies substantially, with a pronounced proprietary-open-source gap. These findings show that advances in visual realism have not yet translated into reliable modeling of scientific and causal dynamics.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!