Long-CODE: 비디오 평가에서 순수한 장기 맥락을 직교적인 차원으로 분리하는 방법
Long-CODE: Isolating Pure Long-Context as an Orthogonal Dimension in Video Evaluation
비디오 생성 모델이 전례 없는 성능을 달성함에 따라, 안정적인 비디오 평가 지표에 대한 수요가 점점 더 중요해지고 있습니다. 기존 지표는 주로 짧은 비디오 평가를 위해 설계되었으며, 프레임 수준의 시각적 품질과 국소적인 시간적 부드러움을 평가하는 데 중점을 둡니다. 그러나 최첨단 비디오 생성 모델이 더 긴 비디오를 생성할 수 있게 되면서, 이러한 지표는 서사적 풍부함과 전역적 인과적 일관성과 같은 중요한 장기적인 특성을 제대로 포착하지 못합니다. 단기적인 시각적 인식과 장기 맥락 속성이 근본적으로 직교하는 차원이라는 점을 인식하여, 우리는 장기 비디오 평가 지표가 단기 비디오 평가와 분리되어야 한다고 주장합니다. 본 논문에서는 장기 비디오 평가를 위한 전용 프레임워크의 엄격한 근거 및 설계를 제시합니다. 먼저, 샷 수준의 변화나 서사적 재배열과 같은 구조적 불일치에 대한 민감성이 부족하여 기존 단기 비디오 지표의 중요한 한계를 드러내는 일련의 장기 비디오 속성 손상 테스트를 소개합니다. 이러한 격차를 해소하기 위해, 샷 동역학을 기반으로 하는 새로운 장기 비디오 지표를 설계했으며, 이는 장기적인 측면을 평가하는 데 매우 효과적입니다. 또한, 장기 비디오 평가를 위한 벤치마킹을 위해 설계된 Long-CODE (Long-Context as an Orthogonal Dimension for video Evaluation)라는 특수 데이터셋을 소개하며, 이 데이터셋은 진정한 장기적인 특성에 대한 인간의 주석을 포함합니다. 광범위한 실험 결과, 제안된 지표가 인간의 판단과 최첨단 수준의 상관관계를 보이는 것으로 나타났습니다. 궁극적으로, 우리의 지표와 벤치마크는 기존의 단기 비디오 표준을 완벽하게 보완하여, 비디오 생성 모델에 대한 포괄적이고 편향되지 않은 평가 패러다임을 구축합니다.
As video generation models achieve unprecedented capabilities, the demand for robust video evaluation metrics becomes increasingly critical. Traditional metrics are intrinsically tailored for short-video evaluation, predominantly assessing frame-level visual quality and localized temporal smoothness. However, as state-of-the-art video generation models scale to generate longer videos, these metrics fail to capture essential long-range characteristics, such as narrative richness and global causal consistency. Recognizing that short-term visual perception and long-context attributes are fundamentally orthogonal dimensions, we argue that long-video metrics should be disentangled from short-video assessments. In this paper, we focus on the rigorous justification and design of a dedicated framework for long-video evaluation. We first introduce a suite of long-video attribute corruption tests, exposing the critical limitations of existing hort-video metrics from their insensitivity to structural inconsistencies, such as shot-level perturbations and narrative shuffling. To bridge this gap, we design a novel long-video metric based on shot dynamics, which is highly sensitive to the long-range testing framework. Furthermore, we introduce Long-CODE (Long-Context as an Orthogonal Dimension for video Evaluation), a specialized dataset designed to benchmark long-video evaluation, with human annotations isolated specifically to genuine long-range characteristics. Extensive experiments show that our proposed metrics achieve state-of-the-art correlation with human judgments. Ultimately, our metric and benchmark seamlessly complement existing short-video standards, establishing a holistic and unbiased evaluation paradigm for video generation models.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.