TimeBlind: 비디오 LLM을 위한 시공간적 구성성 평가 벤치마크
TimeBlind: A Spatio-Temporal Compositionality Benchmark for Video LLMs
세밀한 시공간적 이해는 비디오 추론 및 인공지능 시스템에 필수적입니다. 그러나 다중 모드 대규모 언어 모델(MLLM)이 정적인 의미론을 잘 이해하는 반면, 시간적 역학에 대한 이해는 여전히 취약합니다. 본 연구에서는 시공간적 이해 능력을 진단하기 위한 벤치마크인 TimeBlind을 제안합니다. 인지 과학에 영감을 받아 TimeBlind은 세밀한 시간적 이해를 세 가지 수준으로 분류합니다. 즉, 기본적인 사건 인식, 사건의 특징 분석, 그리고 사건 간의 상호 의존성에 대한 추론입니다. 기존 벤치마크가 사건 인식과 시간적 추론을 혼합하는 것과 달리, TimeBlind은 최소 쌍(minimal-pairs) 패러다임을 활용합니다. 최소 쌍은 동일한 정적인 시각적 내용을 공유하지만, 시간적 구조만 다른 비디오 쌍으로 구성되며, 상호 보완적인 질문을 사용하여 언어적 편향을 제거합니다. 20개 이상의 최첨단 MLLM(예: GPT-5, Gemini 3 Pro)을 600개의 선별된 데이터셋(2400개의 비디오-질문 쌍)으로 평가한 결과, 가장 성능이 좋은 MLLM의 정확도(쌍 내의 두 비디오를 모두 정확하게 구별하는 비율)는 48.2%에 불과하며, 이는 인간의 성능(98.2%)에 훨씬 미치지 못합니다. 이러한 결과는 최첨단 모델조차도 진정한 시간적 논리보다는 정적인 시각적 단서에 크게 의존한다는 것을 보여주며, TimeBlind은 차세대 비디오 이해를 위한 중요한 진단 도구로 자리매김할 것입니다. 데이터셋 및 코드는 https://baiqi-li.github.io/timeblind_project/ 에서 확인할 수 있습니다.
Fine-grained spatio-temporal understanding is essential for video reasoning and embodied AI. Yet, while Multimodal Large Language Models (MLLMs) master static semantics, their grasp of temporal dynamics remains brittle. We present TimeBlind, a diagnostic benchmark for compositional spatio-temporal understanding. Inspired by cognitive science, TimeBlind categorizes fine-grained temporal understanding into three levels: recognizing atomic events, characterizing event properties, and reasoning about event interdependencies. Unlike benchmarks that conflate recognition with temporal reasoning, TimeBlind leverages a minimal-pairs paradigm: video pairs share identical static visual content but differ solely in temporal structure, utilizing complementary questions to neutralize language priors. Evaluating over 20 state-of-the-art MLLMs (e.g., GPT-5, Gemini 3 Pro) on 600 curated instances (2400 video-question pairs), reveals that the Instance Accuracy (correctly distinguishing both videos in a pair) of the best performing MLLM is only 48.2%, far below the human performance (98.2%). These results demonstrate that even frontier models rely heavily on static visual shortcuts rather than genuine temporal logic, positioning TimeBlind as a vital diagnostic tool for next-generation video understanding. Dataset and code are available at https://baiqi-li.github.io/timeblind_project/ .
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.