2601.08620v1 Jan 13, 2026 cs.AI

ViDoRe V3: 복잡한 실제 시나리오에서의 검색 증강 생성에 대한 포괄적 평가

ViDoRe V3: A Comprehensive Evaluation of Retrieval Augmented Generation in Complex Real-World Scenarios

Ant'onio Loison
Ant'onio Loison
Citations: 174
h-index: 6
Quentin Macé
Quentin Macé
Illuin Technology
Citations: 76
h-index: 3
Antoine Edy
Antoine Edy
Citations: 15
h-index: 1
Tom Balough
Tom Balough
Citations: 101
h-index: 4
Manuel Faysse
Manuel Faysse
Meta AI
Citations: 720
h-index: 13
C. Hudelot
C. Hudelot
Citations: 372
h-index: 8
Gautier Viaud
Gautier Viaud
Citations: 442
h-index: 9
G. Moreira
G. Moreira
Citations: 308
h-index: 7
Bo Liu
Bo Liu
Citations: 92
h-index: 4
Victor Xing
Victor Xing
Citations: 78
h-index: 5

검색 증강 생성(RAG) 파이프라인은 단순한 단일 문서 검색을 넘어 시각적 요소(표, 차트, 이미지) 해석, 여러 문서 간의 정보 종합, 정확한 출처 근거 제공과 같은 과제들을 해결해야 합니다. 기존 벤치마크들은 주로 텍스트 데이터나 단일 문서 이해에 집중하거나 검색과 생성을 분리하여 평가함으로써 이러한 복잡성을 제대로 포착하지 못하고 있습니다. 본 연구에서는 시각적으로 풍부한 문서 말뭉치에 대한 다양한 유형의 질의를 특징으로 하는 포괄적인 멀티모달 RAG 벤치마크인 ViDoRe v3를 소개합니다. 이 벤치마크는 다양한 전문 분야에 걸친 10개 데이터셋을 포괄하며, 약 26,000페이지의 문서와 6개 국어로 제공되는 3,099개의 인간 검증 질의로 구성되어 있습니다. 12,000시간에 달하는 인간 주석 작업을 통해 검색 관련성, 바운딩 박스 위치 지정(localization), 검증된 참조 답변에 대한 고품질 주석을 제공합니다. 최신 RAG 파이프라인을 평가한 결과, 시각적 검색기(visual retriever)가 텍스트 검색기보다 우수한 성능을 보였으며, 후기 상호작용(late-interaction) 모델과 텍스트 리랭킹(reranking)이 성능을 크게 향상시키고, 하이브리드 또는 순수 시각적 맥락이 답변 생성 품질을 높이는 것으로 나타났습니다. 그러나 현재 모델들은 여전히 비텍스트 요소, 개방형 질의, 정교한 시각적 근거(fine-grained visual grounding) 처리에 어려움을 겪고 있습니다. 이러한 과제 해결의 진전을 독려하기 위해 본 벤치마크는 https://hf.co/vidore 에서 상업적으로 이용 가능한 라이선스로 공개됩니다.

Original Abstract

Retrieval-Augmented Generation (RAG) pipelines must address challenges beyond simple single-document retrieval, such as interpreting visual elements (tables, charts, images), synthesizing information across documents, and providing accurate source grounding. Existing benchmarks fail to capture this complexity, often focusing on textual data, single-document comprehension, or evaluating retrieval and generation in isolation. We introduce ViDoRe v3, a comprehensive multimodal RAG benchmark featuring multi-type queries over visually rich document corpora. It covers 10 datasets across diverse professional domains, comprising ~26,000 document pages paired with 3,099 human-verified queries, each available in 6 languages. Through 12,000 hours of human annotation effort, we provide high-quality annotations for retrieval relevance, bounding box localization, and verified reference answers. Our evaluation of state-of-the-art RAG pipelines reveals that visual retrievers outperform textual ones, late-interaction models and textual reranking substantially improve performance, and hybrid or purely visual contexts enhance answer generation quality. However, current models still struggle with non-textual elements, open-ended queries, and fine-grained visual grounding. To encourage progress in addressing these challenges, the benchmark is released under a commercially permissive license at https://hf.co/vidore.

16 Citations
3 Influential
6.5 Altmetric
54.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!