LakeQuest: 데이터 레이크 기반 질의응답을 위한 세 가지 영역 통합 벤치마크
LakeQuest: A Three-Domain Benchmark for Grounded Question Answering across Data Lakes
최신 질의응답(QA) 시스템은 체계화된 데이터셋에서 뛰어난 성능을 보이지만, 실제 세상의 지식은 종종 이러한 형태로 정리되어 있지 않습니다. 기업 및 과학 분야의 데이터 레이크에 대한 질문에 답하려면 시스템이 이질적이고 구조가 약한 테이블, 텍스트, 연결된 메타데이터 컬렉션을 탐색해야 합니다. 현재 벤치마크는 이러한 노이즈가 많은 탐색 과정을 간과하여 전체적인 성능을 평가하지 못합니다. 이러한 격차를 해소하기 위해 우리는 현실적인 데이터 레이크에 대한 엔드 투 엔드 검색 및 합성 파이프라인의 성능을 평가하도록 설계된, 사람이 검증한 9,846개의 질의응답 쌍으로 구성된 벤치마크인 LakeQuest를 소개합니다. LakeQuest는 세 가지 다양한 영역(AI/ML 메타데이터, 소매 금융, 다중 모달 생물 의약 정보)을 포괄하며, 각 질문에 대해 정확하고 모드별로 인식하는 증거 지표를 제공합니다. LakeQuest는 데이터 출처 검색과 모달 간 합성 과정을 분리하여 현대 QA 시스템의 중요한 오류 유형을 드러냅니다. 표준 Retrieval-Augmented Generation (RAG) 및 에이전트 기반 도구 사용 방법을 포함한 기본 평가 결과는 고품질 검색만으로는 정확한 추론을 보장할 수 없음을 보여줍니다. 시스템은 메타데이터 그래프에서의 관계 연결, 은행 장부 내 정책 적용, 생물 의약 분야의 합동 테이블 QA에서 일관되게 어려움을 겪으며, 이는 향후 에이전트 기반 QA 시스템에서 강력한 탐색 및 정확한 파일 간 통합 메커니즘의 필요성을 강조합니다.
While modern question answering (QA) systems excel on clean, schema-aligned corpora, real-world knowledge is rarely so neatly packaged. Answering questions over enterprise and scientific data lakes requires systems to navigate heterogeneous, weakly structured collections of tables, passages, and linked metadata. Current benchmarks abstract away this noisy discovery process, failing to evaluate end-to-end performance. To bridge this gap, we introduce LakeQuest, a human-validated benchmark of 9,846 QA pairs designed to evaluate the end-to-end retrieve-and-synthesize pipeline over realistic data lakes. LakeQuest spans three diverse domains (AI/ML metadata, retail banking, and multimodal biomedical drug information) and pairs every question with exact, modality-aware evidence pointers. By isolating source discovery from cross-modal synthesis, LakeQuest exposes critical failure modes in modern QA systems. Our baseline evaluations, including standard Retrieval-Augmented Generation (RAG) and agentic tool-use methods, reveal that high-quality retrieval does not guarantee correct reasoning. Systems consistently struggle with relation chaining in metadata graphs, policy grounding in bank ledgers, and joint tabular QA in biomedical contexts, highlighting the need for robust discovery and faithful cross-file composition mechanisms in future agentic QA systems.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.