물리적 질문 장면 그래프(Physics Question Scene Graph): 텍스트-비디오 생성에서의 물리적 타당성 미세 평가
Physics Question Scene Graph: Fine-grained Evaluation of Physical Plausibility in Text-to-Video Generation
비디오 생성 모델은 점점 더 사실적인 비디오를 생성하는 능력을 갖추고 있지만, 여전히 기본적인 물리 법칙을 따르는 비디오를 생성하는 데 어려움을 겪습니다. 이러한 문제점에 더해, 비디오 내의 물리 법칙 위반 사항을 정확하게 파악하고 규정할 수 있는 신뢰성 있는 미세 평가 방법이 부족합니다. 본 연구에서는 Physics Question Scene Graph (PQSG)라는 계층적 질문 기반 평가 파이프라인을 제안하여 이러한 문제를 해결하고자 합니다. PQSG는 vision-language model (VLM)에 의해 생성된 그래프 기반의 계층적인 질문들을 사용하여, 객체, 동작 및 물리 법칙 준수 측면에서 생성된 비디오가 주어진 프롬프트와 얼마나 일치하는지를 평가합니다. PQSG는 질문을 그래프로 표현함으로써 질문 간의 논리적 의존성을 도입하여 각 질문이 맥락적으로 유효하도록 합니다. 또한, PQSG는 비디오의 어떤 특징이 물리적 타당성 제약 조건을 위반하는지에 대한 미세한 평가를 제공합니다. 본 연구에서는 FinePhyEval이라는 데이터셋을 구축하여, 다양한 최첨단 비디오 생성 모델(Sora 2, Veo 3, Wan 2.1)에서 생성된 비디오와 물리 기반 프롬프트를 함께 제공하고, 각 비디오에 대해 사람이 여러 범주로 주석을 달았습니다. FinePhyEval 데이터셋을 사용하여 PQSG의 미세한 점수와 인간 평가 간의 상관관계를 측정하였으며, 이전 연구보다 높은 전체적인 상관관계를 보였습니다. 또한, PQSG는 폐쇄형 모델들이 Wan 2.1 모델보다 물리적 현실성 측면에서 더 높은 순위를 차지하는 것을 확인했습니다. 마지막으로, FinePhyEval에 제공된 주석이 하위 작업 평가에도 사용될 수 있음을 보여줍니다. 즉, 두 개의 강력한 VLMs를 사용하여 질문 생성 및 답변 성능을 벤치마킹한 결과, 모델들이 인간과 유사한 질문을 생성할 수는 있지만, 여전히 질문 답변 능력에서는 인간의 수준에 미치지 못한다는 것을 확인했습니다.
Video generation models are increasingly capable of producing realistic videos, but they still struggle to generate videos that follow basic physical laws. Compounding this is a lack of reliable granular evaluation methods for localizing and specifying physical law violations in videos. We address this by introducing Physics Question Scene Graph (PQSG), a hierarchical question-based evaluation pipeline. PQSG evaluates generated videos by checking their faithfulness to a prompt across objects, actions, and adherence to physical laws using a graph-based hierarchy of questions generated by a vision-language model (VLM), guided by high-quality in-context examples. By representing questions as a graph, PQSG introduces logical dependencies within questions, ensuring that each query is contextually valid. Moreover, PQSG provides granular assessments of which qualities of the video violate physical plausibility constraints. We validate PQSG by creating FinePhyEval, a dataset with physics-based prompts and corresponding generated videos from diverse state-of-the-art video generation models (Sora 2, Veo 3, and Wan 2.1), with each video annotated across multiple categories by humans. Using FinePhyEval, we measure the correlation between PQSG's fine-grained scores and human judgments, showing higher overall correlations than prior work. We also find that PQSG ranks closed-source models higher than Wan 2.1 on physical realism. Lastly, we show that the annotations we provide in FinePhyEval can also be used for subtask evaluation: we benchmark two strong VLMs on generating and answering questions, finding that while models can create human-like questions, they still fall short of human performance in answering them.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.