2607.06482v1 Jul 07, 2026 cs.CL

현장 데이터 분석: 실제 데이터 복잡성을 고려한 대규모 언어 모델 성능 평가

Data Analysis in the Wild: Benchmarking Large Language Models Against Real-World Data Complexities

Lei Liu
Lei Liu
Citations: 5
h-index: 1
So Hasegawa
So Hasegawa
Citations: 5
h-index: 1
Shailaja Keyur Sampat
Shailaja Keyur Sampat
Arizona State University
Citations: 1,423
h-index: 9
Wei-Peng Chen
Wei-Peng Chen
Citations: 12
h-index: 3

대규모 언어 모델(LLM)의 데이터 분석 능력을 평가하는 현재의 벤치마크는 종종 현실 세계의 환경을 제대로 반영하지 못합니다. 이들은 주로 작은 테이블에서 사실 정보를 검색하는 데 초점을 맞추고, 대규모 다중 테이블 데이터 세트, 외부 지식 통합 및 탐색적 인사이트 발견과 같은 과제를 간과합니다. 본 연구에서는 정부 개방형 데이터를 기반으로 설계되어 실제 시나리오에서 LLM을 평가하기 위한 벤치마크인 DataGovBench를 소개합니다. 이 벤치마크는 복잡한 분해 가능한 질문에 답하고 텍스트 답변 또는 시각화를 생성하는 Table QA와, 탐색적 데이터 분석을 통해 전문가 수준의 발견 결과를 생성하는 모델의 능력을 평가하는 Table Insight라는 두 가지 과제를 포함합니다. 최첨단 LLM을 사용하여 수행한 광범위한 실험 결과, 에이전트 프레임워크를 사용했는지 여부에 관계없이 모든 과제에서 상당한 성능 격차가 나타났습니다. 이러한 결과는 현재 LLM 기반 시스템이 실제 데이터 분석의 요구 사항을 충족하기에 훨씬 부족하다는 것을 시사합니다. DataGovBench는 분석적 질문에 답변하고 데이터로부터 인사이트를 발견할 수 있는 LLM 연구 발전에 도움이 되는 도전적인 벤치마크를 제공합니다. 코드 및 샘플 데이터는 https://github.com/SoHasegawa/datagovbench 에서 확인할 수 있습니다.

Original Abstract

Current benchmarks for evaluating Large Language Models (LLMs) in data analysis often fail to reflect real-world settings. They typically focus on fact retrieval from small tables and overlook the challenges of large multi-tabular datasets, external knowledge integration, and exploratory insight discovery. We introduce DataGovBench, a benchmark derived from governmental open data designed to evaluate LLMs in practical scenarios. The benchmark includes two tasks: Table QA that requires solving complex decomposable questions and producing textual answers or visualizations, and Table Insight that evaluates the ability of models to generate expert-level findings through exploratory data analysis. Comprehensive experiments with state-of-the-art LLMs, both with and without agentic frameworks, reveal significant performance gaps across both tasks. These results suggest that current LLM-based systems remain far from satisfying the demands of real-world data analytics. DataGovBench provides a challenging benchmark for advancing research on LLMs capable of both answering analytical queries and discovering insights from data. Code and sample data are available at https://github.com/SoHasegawa/datagovbench.

0 Citations
0 Influential
24.5 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!