VGA-Bench: 비디오 심미성과 생성 품질 평가를 위한 통합 벤치마크 및 다중 모델 프레임워크
VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation
인공지능 기반 비디오 생성 기술의 빠른 발전은 기존의 생성 품질 지표를 넘어 심미적 매력을 포괄하는 종합적인 평가 프레임워크의 필요성을 강조하고 있습니다. 그러나 기존 벤치마크는 기술적인 정확성에 주로 집중되어 있으며, 특히 인간의 인지적 및 예술적 판단과 관련된 전반적인 평가에는 상당한 격차가 존재합니다. 이러한 한계를 극복하기 위해, 본 논문에서는 비디오 생성 품질과 심미적 품질을 동시에 평가하기 위한 통합 벤치마크인 VGA-Bench를 제안합니다. VGA-Bench는 '심미적 품질', '심미적 태깅', '생성 품질'이라는 세 가지 주요 분류 체계로 구성되며, 각 분류는 체계적인 평가를 가능하게 하기 위해 여러 세부적인 하위 차원으로 구성됩니다. 이러한 분류 체계를 바탕으로, 1,016개의 다양한 프롬프트를 설계하고, 12개의 비디오 생성 모델을 사용하여 60,000개 이상의 비디오를 생성하여 콘텐츠, 스타일, 그리고 생성 과정에서 발생하는 문제점 등 다양한 측면을 포괄하는 대규모 데이터셋을 구축했습니다. 확장 가능하고 자동화된 평가를 위해, 데이터셋의 일부에 대해 인간의 판단을 반영한 레이블링을 수행하고, 심미적 품질 예측을 위한 VAQA-Net, 자동 심미적 태깅을 위한 VTag-Net, 그리고 생성 품질 및 기본적인 품질 속성 평가를 위한 VGQA-Net이라는 세 가지 다중 작업 신경망 평가 모델을 개발했습니다. 광범위한 실험 결과, 제안하는 모델들은 인간의 판단과 높은 신뢰도를 보이는 것으로 나타났으며, 정확성과 효율성을 모두 제공합니다. VGA-Bench는 AIGC 평가 연구를 촉진하기 위해 공개 벤치마크로 제공되며, 콘텐츠 검열, 모델 디버깅, 그리고 생성 모델 최적화에 활용될 수 있습니다.
The rapid advancement of AIGC-based video generation has underscored the critical need for comprehensive evaluation frameworks that go beyond traditional generation quality metrics to encompass aesthetic appeal. However, existing benchmarks remain largely focused on technical fidelity, leaving a significant gap in holistic assessment-particularly with respect to perceptual and artistic qualities. To address this limitation, we introduce VGA-Bench, a unified benchmark for joint evaluation of video generation quality and aesthetic quality. VGA-Bench is built upon a principled three-tier taxonomy: Aesthetic Quality, Aesthetic Tagging, and Generation Quality, each decomposed into multiple fine-grained sub-dimensions to enable systematic assessment. Guided by this taxonomy, we design 1,016 diverse prompts and generate a large-scale dataset of over 60,000 videos using 12 video generation models, ensuring broad coverage across content, style, and artifacts. To enable scalable and automated evaluation, we annotate a subset of the dataset via human labeling and develop three dedicated multi-task neural assessors: VAQA-Net for aesthetic quality prediction, VTag-Net for automatic aesthetic tagging, and VGQA-Net for generation and basic quality attributes. Extensive experiments demonstrate that our models achieve reliable alignment with human judgments, offering both accuracy and efficiency. We release VGA-Bench as a public benchmark to foster research in AIGC evaluation, with applications in content moderation, model debugging, and generative model optimization.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.