2602.00288v3 Jan 30, 2026 cs.CV

TimeBlind: 비디오 LLM을 위한 시공간적 구성성 평가 벤치마크

TimeBlind: A Spatio-Temporal Compositionality Benchmark for Video LLMs

Gedas Bertasius
Gedas Bertasius
Citations: 7,648
h-index: 29
Baiqi Li
Baiqi Li
Citations: 747
h-index: 5
Kangyi Zhao
Kangyi Zhao
Citations: 19
h-index: 3
Ce Zhang
Ce Zhang
Citations: 42
h-index: 3
Chancharik Mitra
Chancharik Mitra
Citations: 380
h-index: 6
Jean de Dieu Nyandwi
Jean de Dieu Nyandwi
Citations: 177
h-index: 4

세밀한 시공간적 이해는 비디오 추론 및 인공지능 시스템에 필수적입니다. 그러나 다중 모드 대규모 언어 모델(MLLM)이 정적인 의미론을 잘 이해하는 반면, 시간적 역학에 대한 이해는 여전히 취약합니다. 본 연구에서는 시공간적 이해 능력을 진단하기 위한 벤치마크인 TimeBlind을 제안합니다. 인지 과학에 영감을 받아 TimeBlind은 세밀한 시간적 이해를 세 가지 수준으로 분류합니다. 즉, 기본적인 사건 인식, 사건의 특징 분석, 그리고 사건 간의 상호 의존성에 대한 추론입니다. 기존 벤치마크가 사건 인식과 시간적 추론을 혼합하는 것과 달리, TimeBlind은 최소 쌍(minimal-pairs) 패러다임을 활용합니다. 최소 쌍은 동일한 정적인 시각적 내용을 공유하지만, 시간적 구조만 다른 비디오 쌍으로 구성되며, 상호 보완적인 질문을 사용하여 언어적 편향을 제거합니다. 20개 이상의 최첨단 MLLM(예: GPT-5, Gemini 3 Pro)을 600개의 선별된 데이터셋(2400개의 비디오-질문 쌍)으로 평가한 결과, 가장 성능이 좋은 MLLM의 정확도(쌍 내의 두 비디오를 모두 정확하게 구별하는 비율)는 48.2%에 불과하며, 이는 인간의 성능(98.2%)에 훨씬 미치지 못합니다. 이러한 결과는 최첨단 모델조차도 진정한 시간적 논리보다는 정적인 시각적 단서에 크게 의존한다는 것을 보여주며, TimeBlind은 차세대 비디오 이해를 위한 중요한 진단 도구로 자리매김할 것입니다. 데이터셋 및 코드는 https://baiqi-li.github.io/timeblind_project/ 에서 확인할 수 있습니다.

Original Abstract

Fine-grained spatio-temporal understanding is essential for video reasoning and embodied AI. Yet, while Multimodal Large Language Models (MLLMs) master static semantics, their grasp of temporal dynamics remains brittle. We present TimeBlind, a diagnostic benchmark for compositional spatio-temporal understanding. Inspired by cognitive science, TimeBlind categorizes fine-grained temporal understanding into three levels: recognizing atomic events, characterizing event properties, and reasoning about event interdependencies. Unlike benchmarks that conflate recognition with temporal reasoning, TimeBlind leverages a minimal-pairs paradigm: video pairs share identical static visual content but differ solely in temporal structure, utilizing complementary questions to neutralize language priors. Evaluating over 20 state-of-the-art MLLMs (e.g., GPT-5, Gemini 3 Pro) on 600 curated instances (2400 video-question pairs), reveals that the Instance Accuracy (correctly distinguishing both videos in a pair) of the best performing MLLM is only 48.2%, far below the human performance (98.2%). These results demonstrate that even frontier models rely heavily on static visual shortcuts rather than genuine temporal logic, positioning TimeBlind as a vital diagnostic tool for next-generation video understanding. Dataset and code are available at https://baiqi-li.github.io/timeblind_project/ .

6 Citations
2 Influential
14.5 Altmetric
82.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!