2604.08008v1 Apr 09, 2026 cs.CV

SearchAD: 자율 주행을 위한 대규모 희귀 이미지 검색 데이터셋

SearchAD: Large-Scale Rare Image Retrieval Dataset for Autonomous Driving

Marius Cordts
Marius Cordts
Citations: 14
h-index: 2
Markus Enzweiler
Markus Enzweiler
Citations: 8
h-index: 2
Felix Embacher
Felix Embacher
Citations: 1
h-index: 1
J. Uhrig
J. Uhrig
Citations: 2,080
h-index: 7

견고한 자율 주행(AD) 시스템을 구축하기 위해서는 대규모 데이터셋에서 희귀하고 안전에 중요한 주행 시나리오를 검색하는 것이 필수적입니다. 데이터셋 크기가 계속 증가함에 따라, 주요 과제는 더 많은 데이터를 수집하는 것이 아니라 가장 관련성이 높은 샘플을 효율적으로 식별하는 것으로 전환됩니다. 본 논문에서는 11개의 기존 데이터셋에서 추출한 42만 3천 프레임을 포함하는 대규모 희귀 이미지 검색 데이터셋인 SearchAD를 소개합니다. SearchAD는 90개의 희귀 카테고리에 대한 51만 3천 개 이상의 바운딩 박스를 포함하는 고품질의 수동 주석을 제공합니다. SearchAD는 특히 전체 데이터셋에서 50번 미만으로 나타나는 극히 희귀한 클래스를 찾는 '바늘 찾기' 문제를 목표로 합니다. 기존 벤치마크와 달리, SearchAD는 인스턴스 수준 검색에 초점을 맞추기보다는 명확하게 정의된 데이터 분할을 통해 텍스트-이미지 및 이미지-이미지 검색, Few-shot 학습, 그리고 다중 모드 검색 모델의 미세 조정을 가능하게 하는 의미론적 이미지 검색을 강조합니다. 종합적인 평가 결과, 텍스트 기반 방법이 더 강력한 의미론적 기반을 갖추고 있어 이미지 기반 방법보다 성능이 우수합니다. 시각적 특징과 언어를 직접적으로 연결하는 모델이 가장 우수한 제로샷 결과를 달성했지만, 우리의 미세 조정 기준 모델은 성능을 크게 향상시켰습니다. SearchAD는 공개 벤치마크 서버에 제공된 별도의 테스트 세트를 통해 자율 주행 분야에서 검색 기반 데이터 큐레이션 및 롱테일 인식 연구를 위한 최초의 대규모 데이터셋을 제공합니다: https://iis-esslingen.github.io/searchad/.

Original Abstract

Retrieving rare and safety-critical driving scenarios from large-scale datasets is essential for building robust autonomous driving (AD) systems. As dataset sizes continue to grow, the key challenge shifts from collecting more data to efficiently identifying the most relevant samples. We introduce SearchAD, a large-scale rare image retrieval dataset for AD containing over 423k frames drawn from 11 established datasets. SearchAD provides high-quality manual annotations of more than 513k bounding boxes covering 90 rare categories. It specifically targets the needle-in-a-haystack problem of locating extremely rare classes, with some appearing fewer than 50 times across the entire dataset. Unlike existing benchmarks, which focused on instance-level retrieval, SearchAD emphasizes semantic image retrieval with a well-defined data split, enabling text-to-image and image-to-image retrieval, few-shot learning, and fine-tuning of multi-modal retrieval models. Comprehensive evaluations show that text-based methods outperform image-based ones due to stronger inherent semantic grounding. While models directly aligning spatial visual features with language achieve the best zero-shot results, and our fine-tuning baseline significantly improves performance, absolute retrieval capabilities remain unsatisfactory. With a held-out test set on a public benchmark server, SearchAD establishes the first large-scale dataset for retrieval-driven data curation and long-tail perception research in AD: https://iis-esslingen.github.io/searchad/

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!