2606.05702v1 Jun 04, 2026 cs.AI

시간을 읽는 능력: 비전-언어 모델에서 시간 순서 추론 및 단축 경로 편향 성능 평가

Seeing Time: Benchmarking Chronological Reasoning and Shortcut Biases in Vision-Language Models

Yongcheng Jing
Yongcheng Jing
Citations: 141
h-index: 7
Qixin Zhang
Qixin Zhang
Citations: 12
h-index: 2
Ziqi Xu
Ziqi Xu
Citations: 24
h-index: 3
Qing Qing
Qing Qing
Citations: 4
h-index: 1
Juncheng Hu
Juncheng Hu
Citations: 120
h-index: 6
Renqian Luo
Renqian Luo
Citations: 36
h-index: 3
Hao Zhou
Hao Zhou
Citations: 195
h-index: 5
Caichong Li
Caichong Li
Citations: 0
h-index: 0
Xikun Zhang
Xikun Zhang
Citations: 8
h-index: 1

최근 비전-언어 모델(VLMs)의 발전은 복잡한 시각적 의미를 해석하는 능력을 크게 향상시켰지만, 시간 순서에 대한 추론 능력은 아직 충분히 연구되지 않았습니다. 본 논문에서는 VLM이 이미지 내외부에서 시간 정보를 어떻게 인식하고 추론하는지 평가하기 위해 특별히 설계된 새로운 벤치마크를 소개합니다. 기존의 비디오 기반 벤치마크가 프레임 순서에 초점을 맞추는 것과 달리, 우리는 시간적 판단의 근본적인 논리와 다중 모달 통합 측면을 탐구합니다. 이를 위해 세 가지 특수 데이터셋을 구축했습니다. 첫 번째 데이터셋은 긴 역사적 기간 동안 시각적으로 유사한 객체를 포함하고, 두 번째 데이터셋은 다양한 이벤트 및 객체 유형으로 분류되어 있으며, 세 번째 데이터셋은 시간 민감한 뉴스 텍스트와 이미지를 쌍으로 연결하여 다중 모달 정렬을 수행합니다. 광범위한 실험을 통해 모델이 범주별로 성능 차이를 보이는지 분석하고, 더욱 중요한 점은 모델이 실제 시간적 특징 대신 이미지 색상과 같은 "잘못된 단축 경로"에 의존하는지 여부를 탐구했습니다. 그 결과, VLM은 잠재력을 보여주지만, 종종 진정한 시간적 추론을 우회하기 위해 명암 대비와 같은 피상적인 신호를 활용합니다. 고품질 데이터셋과 엄격한 평가 프레임워크를 제공함으로써, 우리는 현재의 한계를 파악하고 더욱 강력하고 논리적으로 기반한 다중 모달 모델 개발을 위한 진단 도구를 제시합니다. 소스 코드는 https://github.com/LuoRenqiang/ChronoVision 에서 확인할 수 있습니다.

Original Abstract

Recent advancements in Vision-Language Models (VLMs) have significantly enhanced their ability to interpret complex visual semantics, yet their capacity for chronological reasoning remains under-explored. In this paper, we introduce a novel benchmark specifically designed to evaluate how VLMs perceive and reason about chronological information within and across images. Unlike existing video-based benchmarks that focus on frame sequencing, our work delves into the underlying logic of chronological judgment and the expansion toward multimodal integration. To facilitate this, we construct three specialized datasets: one containing visually similar objects spanning long historical durations, another categorized by diverse event and object types, and a third pairing images with time-sensitive news text for cross-modal alignment. Through extensive experiments, we analyze whether models exhibit performance disparities across categories and, crucially, explore whether they rely on ``incorrect shortcuts'', such as image color rather than genuine chronological features. Our results reveal that while VLMs show promise, they frequently exploit superficial cues like grayscale versus color filters to bypass authentic chronological reasoning. By providing these high-quality datasets and a rigorous evaluation framework, we offer a diagnostic tool to identify current limitations and guide the development of more robust, logically grounded multimodal models. The source code is shown in https://github.com/LuoRenqiang/ChronoVision.

0 Citations
0 Influential
23.5 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!