저주파 함정: 비디오 언어 모델이 단순한 이벤트 기록에서 실패하는 이유
The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping
실제 비디오 벤치마크는 광범위한 내용을 제공하지만, 고정된 클립들은 이벤트 발생 횟수, 빈도, 지속 시간 및 시각적 복잡성을 함께 묶어 오류의 원인을 파악하기 어렵게 만듭니다. 기존의 프로그래밍 기반 벤치마크는 더 나은 제어를 제공하지만, 보고된 이벤트를 실행 가능한 기준 사실과 비교하여 평가하는 대신 최종 답변만 평가합니다. 이러한 격차를 해소하기 위해, 우리는 세 가지 통제된 비디오 작업(공을 떨어뜨리는 동작 시 벽과의 접촉, 시각적 깜빡임, 범주형 상태 변화)에서 이벤트 수를 계산하기 위한 추적 기반의 매개변수 프로파일링 방법을 소개합니다. 2,190개의 비디오를 사용하여 렌더링 방식을 고정하고 이벤트 발생 횟수(N)와 빈도(F)를 변경했습니다. 각 비디오에는 기능 면적 추정을 위한 실행 가능한 이벤트 추적 정보와 타임스탬프 수준 평가가 포함되어 있습니다. 우리의 결과는 단계적인 시간 관련 오류를 보여줍니다. 80%의 신뢰성 기준으로, Gemini 3.6 Flash 모델은 지속적인 상태 변화를 최대 12회까지 0.5Hz 및 1.0Hz에서 안정적으로 계산할 수 있지만, 일시적인 깜빡임 이벤트에 대해서는 신뢰할 만한 정확한 계산 결과를 전혀 보여주지 않습니다. 즉, 이벤트 표현 방식이 모델이 초기 증거를 활용하는 방식을 결정하며, 이는 이벤트 발생 횟수와 빈도가 증가함에 따라 더욱 심화됩니다. 높은 이벤트 발생 횟수 및 높은 빈도에서, 최종 계산 값 중 올바른 결과는 0.2%에 불과하며, 모델은 실제 이벤트의 18.1%만을 정확하게 식별합니다. 시각적 정보 접근이 주요 병목 현상인지 확인하기 위해, 샘플링 속도를 높였습니다. 이로 인해 공을 떨어뜨리는 동작의 정확도가 19.6%에서 29.3%로 향상되었지만, 보고된 이벤트 순서가 기준 사실과 일치하는 경우는 3.7%에 불과했습니다. 따라서 추가 프레임은 최종 점수를 높일 수 있지만, 실제 이벤트 복구에는 도움이 되지 않을 수 있습니다. 다양한 프롬프트 전략을 사용해도 유사한 제한적인 개선 효과만 나타났으며, 실제 비디오 평가에서도 낮은 이벤트 발생 횟수에서 성공률이 집중되는 동일한 경향이 관찰되었습니다. 궁극적으로, 추적 기반 프로파일링은 비디오 평가를 집계 정확도 지표에서 벗어나 시간 관련 추론 실패의 세부적인 원인 분석으로 전환하는 데 기여합니다.
Real-world video benchmarks provide broad coverage, but their fixed clips entangle event count, rate, duration, and visual complexity, making failure modes hard to isolate. While existing programmatic benchmarks offer better control, they score only the final answer rather than auditing reported events against executable ground truth. To bridge this gap, we introduce trace-grounded parametric profiling for event counting in three controlled video tasks: bouncing-ball wall contacts, visual blinks, and categorical state transitions. Across 2,190 videos, we vary event count N and frequency F while holding rendering fixed. Each video includes an executable event trace for capability-surface estimation and timestamp-level evaluation. Our results reveal a staged temporal failure. At an 80% reliability threshold, Gemini 3.6 Flash reliably counts persistent state transitions up to 12 events at 0.5 and 1.0 Hz, yet demonstrates no reliable positive-count region for transient blinking events. Thus, event representation dictates whether a model initially accesses evidence -- a limitation that compounds as count and frequency increase. In the high-count, high-frequency regime, only 0.2% of final counts are correct and the model recovers just 18.1% of true events. To test if visual access is the primary bottleneck, we increase sampling rate. Although this boosts Bounce Ball accuracy from 19.6% to 29.3%, the reported sequence agrees with ground truth only 3.7% of the time. Extra frames can therefore inflate final scores without producing faithful event recovery. Different prompting strategies yield similarly limited gains, and real-world video evaluations show the same concentration of success at low event counts. Ultimately, trace-grounded profiling shifts video evaluation from aggregate accuracy metrics to a detailed diagnostic of where temporal reasoning fails.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.