ArtiFact: 대규모 다중 모드 문화유산 데이터셋
ArtiFact: A Large-Scale Multi-Modal Cultural Heritage Dataset
다중 모드 데이터 관리는 데이터 통합, 의미 기반 질의 처리 및 데이터 품질 평가를 포함하는 데이터베이스 분야의 핵심 연구 주제로 부상했습니다. 이러한 관심이 높아짐에도 불구하고, 테이블, 텍스트 및 이미지를 결합한 대규모 실세계 데이터셋은 부족합니다. 본 논문에서는 메트로폴리탄 미술관, 시카고 예술 협회 및 국립 박물관에서 수집된 651,045개의 박물관 기록으로 구성된 다중 모드 문화유산 데이터셋인 ArtiFact를 소개합니다. 우리는 두 가지 후속 작업을 통해 ArtiFact의 유용성을 보여줍니다. 교차 모드 오류 감지를 위해, 130,209개의 기록에 삽입된 7가지 오류 범주로 구성된 큐레이션된 분류 체계를 도입하고, 재료 시대착오 및 시간 변화와 같은 미묘한 도메인별 오류를 안정적으로 탐지하는 것은 여전히 해결해야 할 과제임을 보여줍니다. 의미 기반 질의 처리에 있어서는 현재 시스템이 문화적 근접성, 모호한 객체 유형 및 역사적 맥락에 따른 용어를 포함하는 질의 처리에서 어려움을 겪는다는 것을 보여줍니다. 우리의 결과는 ArtiFact를 다중 모드 데이터 관리 연구를 위한 도전적인 벤치마크로 자리매김합니다.
Multi-modal data management has emerged as a central research topic in the database community, spanning data integration, semantic query processing, and data quality assessment. Despite this growing interest, the community lacks large-scale, real-world datasets combining tables, text, and images. We present ArtiFact, a multi-modal cultural heritage dataset of 651045 museum records collected from the Metropolitan Museum of Art, the Art Institute of Chicago, and the Rijksmuseum. We demonstrate the utility of ArtiFact through two downstream tasks. For cross-modal error detection, we introduce a curated taxonomy of seven error categories injected into 130209 records and show that reliably detecting subtle domain-specific errors such as material anachronisms and temporal shifts remain an open challenge. For semantic query processing, we show that current systems struggle with queries involving cultural proximity, ambiguous object types, and historically contingent terminology. Our results position ArtiFact as a challenging benchmark for multi-modal data management research.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.