비전-언어 모델에서 나타나는 문화적 시대착오와 시간적 추론에 관한 연구
On the Cultural Anachronism and Temporal Reasoning in Vision Language Models
비전-언어 모델(VLM)은 디지털 아카이브부터 교육 플랫폼에 이르기까지 다양한 문화 유산 자료에 점점 더 많이 활용되고 있습니다. 본 연구에서는 이러한 모델들이 역사적 유물을 해석하는 방식에서 나타나는 근본적인 문제를 지적합니다. 우리는 이러한 현상을 '문화적 시대착오'라고 정의하며, 이는 역사적 대상들을 시간적으로 부적절한 개념, 재료 또는 문화적 틀을 사용하여 오해하는 경향을 의미합니다. 이 현상을 정량화하기 위해, 우리는 비전-언어 모델을 위한 시간적 시대착오 벤치마크(TAB-VLM)를 소개합니다. TAB-VLM은 선사 시대부터 현대에 이르기까지 1,600개의 인도 문화 유물에 대한 시간적 추론 능력을 평가하기 위해 설계된 6개의 범주에 걸쳐 600개의 질문으로 구성된 데이터 세트입니다. 최첨단 모델 10개를 체계적으로 평가한 결과, 벤치마크에서 상당한 결함이 드러났으며, 최고 성능 모델(GPT-5.2)조차도 58.7%의 전체 정확도에 그쳤습니다. 다양한 아키텍처와 규모에서도 성능 격차가 지속되는 것으로 나타나, 이는 모델 크기와 관계없이 시각 AI 시스템에서 문화적 시대착오가 중요한 제한 요소임을 시사합니다. 이러한 연구 결과는 현재 VLM의 기능과 문화 유산 자료, 특히 훈련 데이터에 충분히 반영되지 않은 비서구 시각 문화의 정확한 해석에 필요한 요구 사항 간의 격차를 보여줍니다. 본 벤치마크는 역사적 유물과 상호 작용하는 다중 모드 AI 시스템에서 시간적 인지 능력을 향상시키는 데 기여할 것입니다. 데이터 세트와 코드는 프로젝트 페이지에서 확인할 수 있습니다.
Vision-Language Models (VLMs) are increasingly applied to cultural heritage materials, from digital archives to educational platforms. This work identifies a fundamental issue in how these models interpret historical artifacts. We define this phenomenon as cultural anachronism, the tendency to misinterpret historical objects using temporally inappropriate concepts, materials, or cultural frameworks. To quantify this phenomenon, we introduce the Temporal Anachronism Benchmark for Vision-Language Models (TAB-VLM), a dataset of 600 questions across six categories, designed to evaluate temporal reasoning on 1,600 Indian cultural artifacts spanning prehistoric to modern periods. Systematic evaluations of ten state-of-the-art models reveal significant deficiencies on our benchmark, and even the best model (GPT-5.2) achieves only 58.7% overall accuracy. The performance gap persists across varying architectures and scales, suggesting that cultural anachronism represents a significant limitation in visual AI systems, regardless of model size. These findings highlight the disparity between current VLM capabilities and the requirements for accurately interpreting cultural heritage materials, particularly for non-Western visual cultures underrepresented in training data. Our benchmark provides a foundation for enhancing temporal cognition in multimodal AI systems that interact with historical artifacts. The dataset and code are available in our project page.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.