2606.10460v1 Jun 09, 2026 cs.CL

LakeQA: 백만 규모 데이터 레이크 기반 탐색적 질의응답 벤치마크

LakeQA: An Exploratory QA Benchmark over a Million-Scale Data Lake

Yusen Zhang
Yusen Zhang
Citations: 2
h-index: 1
Eugene Wu
Eugene Wu
Citations: 8
h-index: 2
Eden Wu
Eden Wu
Citations: 60
h-index: 3
Juliana Freire
Juliana Freire
Citations: 78
h-index: 4
Haonan Wang
Haonan Wang
National University of Singapore
Citations: 226
h-index: 7
Jiaxiang Liu
Jiaxiang Liu
Citations: 23
h-index: 3
Yurong Liu
Yurong Liu
Citations: 120
h-index: 5
A. Wijaya
A. Wijaya
Citations: 0
h-index: 0
Tianle Zhou
Tianle Zhou
Citations: 7
h-index: 2
Yijia Chen
Yijia Chen
Citations: 2
h-index: 1
Wanting You
Wanting You
Citations: 2
h-index: 1
Reya Vir
Reya Vir
Citations: 10
h-index: 2
Daniela Pinto
Daniela Pinto
Citations: 0
h-index: 0
Grace Fan
Grace Fan
NYU
Citations: 396
h-index: 8

최근 대규모 언어 모델(LLM)은 명시적으로 제공되거나 쉽게 검색될 수 있는 증거를 기반으로 하는 질의응답(QA) 분야에서 빠른 발전을 보여왔습니다. 반면, 실제 질문은 종종 정확한 증거 문서와 함께 제시되지 않습니다. 유용한 증거는 방대한 데이터 레이크에 존재하며, 이에 대한 검색이 답변을 위한 필수적인 과정입니다. 하지만 대규모 데이터 레이크를 대상으로 검색과 추론 능력을 동시에 요구하는 종합적인 벤치마크가 부족합니다. 이러한 문제를 해결하기 위해, 본 논문에서는 데이터 레이크 기반의 검색 중심 질의응답을 위한 종합적인 벤치마크인 LakeQA를 소개합니다. LakeQA는 위키피디아 및 공개 정부 데이터를 포함한 약 9.5TB 규모의 다양한 텍스트 자원을 활용하여 구축되었으며, 정형 및 비정형 데이터를 모두 포함합니다. 작업의 품질을 보장하기 위해 각 샘플은 최소 한 명 이상의 박사 학위 소지자가 검토했습니다. 각 작업은 장기적인 다단계 추론을 요구하며, 에이전트는 올바른 문서를 검색하고 여러 출처에서 증거를 조합하여 답변을 생성해야 합니다. seven개의 최첨단 LLM에 대한 실험 결과는 LakeQA가 어려운 벤치마크임을 보여줍니다. 예를 들어, GPT-5.2는 LakeQA에서 정확히 일치하는 점수가 18.37%에 불과했습니다. 전반적으로, LakeQA는 현대 데이터 레이크에서 데이터를 검색하고 분석할 수 있는 LLM 에이전트를 개발하기 위한 현실적인 테스트 환경을 제공합니다.

Original Abstract

Recent large language models (LLMs) have shown rapid progress in reading-based question answering (QA), where evidence is explicitly provided or can be trivially retrieved. In contrast, real-world questions are often not paired with accurate evidence documents. The useful evidence resides in massive data lakes, making search a prerequisite for answering. However, there is a lack of comprehensive benchmarks that require both searching and reasoning over large data lakes. To this end, we introduce LakeQA, a comprehensive benchmark for search-centric question answering over data lakes that jointly emphasizes searching and reasoning capabilities. LakeQA is built on a heterogeneous collection of approximately 9.5 TB of text resources from Wikipedia and open-source government data, spanning structured and unstructured data. To ensure task quality, each sample is annotated by at least one Ph.D.-level expert. Each task requires long-horizon multi-hop reasoning with implicit intermediate steps: agents need to discover the correct documents and then compose evidence across sources to produce the answer. Experimental results on seven frontier LLMs demonstrate that LakeQA is challenging. For instance, GPT-5.2 achieves only an exact-match score of 18.37% on LakeQA. Overall, LakeQA provides a realistic testbed for developing LLM agents that can both find and analyze data in modern data lakes.

1 Citations
0 Influential
4 Altmetric
21.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!