LakeQA: 백만 규모 데이터 레이크 기반 탐색적 질의응답 벤치마크
LakeQA: An Exploratory QA Benchmark over a Million-Scale Data Lake
최근 대규모 언어 모델(LLM)은 명시적으로 제공되거나 쉽게 검색될 수 있는 증거를 기반으로 하는 질의응답(QA) 분야에서 빠른 발전을 보여왔습니다. 반면, 실제 질문은 종종 정확한 증거 문서와 함께 제시되지 않습니다. 유용한 증거는 방대한 데이터 레이크에 존재하며, 이에 대한 검색이 답변을 위한 필수적인 과정입니다. 하지만 대규모 데이터 레이크를 대상으로 검색과 추론 능력을 동시에 요구하는 종합적인 벤치마크가 부족합니다. 이러한 문제를 해결하기 위해, 본 논문에서는 데이터 레이크 기반의 검색 중심 질의응답을 위한 종합적인 벤치마크인 LakeQA를 소개합니다. LakeQA는 위키피디아 및 공개 정부 데이터를 포함한 약 9.5TB 규모의 다양한 텍스트 자원을 활용하여 구축되었으며, 정형 및 비정형 데이터를 모두 포함합니다. 작업의 품질을 보장하기 위해 각 샘플은 최소 한 명 이상의 박사 학위 소지자가 검토했습니다. 각 작업은 장기적인 다단계 추론을 요구하며, 에이전트는 올바른 문서를 검색하고 여러 출처에서 증거를 조합하여 답변을 생성해야 합니다. seven개의 최첨단 LLM에 대한 실험 결과는 LakeQA가 어려운 벤치마크임을 보여줍니다. 예를 들어, GPT-5.2는 LakeQA에서 정확히 일치하는 점수가 18.37%에 불과했습니다. 전반적으로, LakeQA는 현대 데이터 레이크에서 데이터를 검색하고 분석할 수 있는 LLM 에이전트를 개발하기 위한 현실적인 테스트 환경을 제공합니다.
Recent large language models (LLMs) have shown rapid progress in reading-based question answering (QA), where evidence is explicitly provided or can be trivially retrieved. In contrast, real-world questions are often not paired with accurate evidence documents. The useful evidence resides in massive data lakes, making search a prerequisite for answering. However, there is a lack of comprehensive benchmarks that require both searching and reasoning over large data lakes. To this end, we introduce LakeQA, a comprehensive benchmark for search-centric question answering over data lakes that jointly emphasizes searching and reasoning capabilities. LakeQA is built on a heterogeneous collection of approximately 9.5 TB of text resources from Wikipedia and open-source government data, spanning structured and unstructured data. To ensure task quality, each sample is annotated by at least one Ph.D.-level expert. Each task requires long-horizon multi-hop reasoning with implicit intermediate steps: agents need to discover the correct documents and then compose evidence across sources to produce the answer. Experimental results on seven frontier LLMs demonstrate that LakeQA is challenging. For instance, GPT-5.2 achieves only an exact-match score of 18.37% on LakeQA. Overall, LakeQA provides a realistic testbed for developing LLM agents that can both find and analyze data in modern data lakes.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.