OmniEdit-Bench: 지시 기반 비디오 편집을 위한 종합적인 성능 평가 기준
OmniEdit-Bench: A Comprehensive Benchmark for Instruction-based Video Editing
지시 기반 비디오 편집(IVE)은 광범위한 응용 가능성을 가진 새로운 분야이지만, 현재 모델의 성능을 평가하는 것은 어려운 과제입니다. 기존의 성능 평가 기준은 이미지 편집에서 비롯된 제한적인 작업 범위로 인해 비디오 고유의 특성을 간과하고, 또한 지시 사항 준수 여부를 제대로 측정하지 못하여 원본 비디오의 강한 시각적 특징으로 인해 잘못된 편집 결과가 높은 점수를 받을 수 있다는 문제가 있습니다. 이러한 문제점을 해결하기 위해, 우리는 IVE를 위한 종합적이고 체계적인 성능 평가 기준을 제시합니다. 저희의 성능 평가 기준은 공간, 시간, 오디오 및 참조 기반 편집 등 다양한 비디오 관련 요소를 포함하여 기존의 프레임 단위 평가 방식을 넘어선 포괄적인 작업을 수행합니다. 또한 명시적인 지시와 암묵적인 지시를 구분하고, 실제 요구 사항을 더 잘 반영하기 위해 추론 기반 시나리오를 포함했습니다. 더욱이 저희는 정확성, 보존성, 현실감 및 일관성을 기준으로 편집 품질을 평가하는 평가 프레임워크를 제안합니다. 이 프레임워크는 인간의 판단과 최첨단 시각-언어 모델을 모두 사용합니다. 지시 사항 준수 여부를 강조하기 위해, 저희는 정확도에 따라 다른 점수를 조정하는 페널티 메커니즘을 도입하여 시각적으로는 타당해 보이지만 실제로는 잘못된 편집 결과가 과대 평가되는 것을 방지합니다. 대표적인 오픈 소스 및 상용 모델에 대한 광범위한 실험 결과, 현재의 IVE 모델은 여전히 만족스러운 수준에 미치지 못한다는 것을 보여줍니다. OmniEdit-Bench는 지시 기반 비디오 편집을 평가하기 위한 종합적이고 신뢰할 수 있는 테스트 환경을 제공하며, 향후 연구 방향에 대한 통찰력을 제시합니다.
Instruction-based video editing (IVE) is an emerging field with broad applications, yet evaluating editing models remains challenging. Existing benchmarks suffer from two major limitations: limited task coverage inherited from image editing, which overlooks video-specific dimensions, and inadequate metrics that fail to measure instruction fidelity, allowing incorrect edits to receive high scores due to strong visual priors from the original video. To address these issues, we introduce a comprehensive and structured benchmark for IVE. Our benchmark decomposes editing tasks into multiple video-specific dimensions, including spatial, temporal, audio, and reference-based editing, extending beyond conventional frame-level evaluation. It also distinguishes explicit and implicit instructions and incorporates reasoning-based scenarios to better reflect real-world requirements. Furthermore, we propose an evaluation framework that assesses editing quality from four complementary dimensions: accuracy, preservation, realism, and consistency, using both human judgments and state-of-the-art vision-language models. To emphasize instruction fidelity, we introduce an accuracy-aware penalty mechanism that conditions other scores on accuracy, preventing visually plausible but incorrect edits from receiving inflated evaluations. Extensive experiments on representative open-source and commercial models show that current IVE models remain far from satisfactory. OmniEdit-Bench provides a comprehensive and reliable testbed for evaluating instruction-based video editing and offers insights into future research directions.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.