2607.06229v1 Jul 07, 2026 cs.CL

Spider 2.0-AIFunc: 실제 데이터 기반 텍스트-SQL 시스템을 AI 내장 SQL 워크플로우로 확장

Spider 2.0-AIFunc: Extending Real-World Text-to-SQL to AI-Native SQL Workflows

Yuxiong He
Yuxiong He
Citations: 1,644
h-index: 19
Julian McAuley
Julian McAuley
Citations: 43
h-index: 3
Zhewei Yao
Zhewei Yao
Citations: 124
h-index: 5
Jixuan Chen
Jixuan Chen
Citations: 1,374
h-index: 9
N. Kuang
N. Kuang
Citations: 31
h-index: 4
Tianyang Liu
Tianyang Liu
UC San Diego
Citations: 1,518
h-index: 9
Canwen Xu
Canwen Xu
Citations: 20
h-index: 2
Fangyu Lei
Fangyu Lei
Citations: 339
h-index: 3
Tao Yu
Tao Yu
Citations: 1,629
h-index: 9

주요 클라우드 데이터 플랫폼은 이제 대규모 언어 모델 기능을 기본 SQL 함수로 제공하여 분석가들이 일반적인 SQL 쿼리 내에서 분류, 필터링, 감성 분석, 추출, 유사 검색 및 집계 작업을 수행할 수 있도록 합니다. 그러나 기존의 텍스트-SQL 벤치마크는 전통적인 SQL만 평가하며, 모델이 이러한 AI 내장 SQL을 생성할 수 있는지에 대한 정보를 제공하지 않습니다. 본 논문에서는 Snowflake 플랫폼에서 제공되는 6가지 유형의 AI 함수를 포함하는 125개의 실제 데이터베이스에 걸쳐 검증된 465개의 인스턴스로 구성된 벤치마크인 Spider 2.0-AIFunc를 소개합니다. 기존의 엔터프라이즈 텍스트-SQL 벤치마크를 기반으로, 저희는 에이전트 기반 파이프라인을 사용하여 원본 작업을 AI 내장 형태로 재구성하고, 동시에 대상 쿼리를 변환하며, 자연어 지침을 개선하여 의도된 AI 내장 솔루션을 명확하게 하고 모호성을 줄입니다. 모든 인스턴스는 출시 전에 시간적으로 분리된 구간에서 다중 라운드 반복 실행 프로토콜을 거쳐 결과의 안정성을 확인합니다. 최첨단 언어 모델 10개를 평가한 결과, 가장 성능이 뛰어난 독점 모델은 67~70%의 실행 정확도를 달성하는 반면, 최고의 오픈 소스 모델은 58.1%의 정확도를 보였습니다. 이러한 성능 차이는 주로 술어 지정 오류, 스키마 연결 및 AI 함수 매개변수 설정에서 비롯됩니다. 기존 텍스트-SQL 문제 해결을 위해 설계된 에이전트 프레임워크(예: 스키마 검색 및 관련 테이블 선택)는 AI 내장 SQL에 효과적으로 적용되지 않습니다. 최소한의 에이전트 구성이 일관되게 더 복잡한 대안보다 우수한 성능을 보이는 것으로 나타났으며, 이는 이러한 프레임워크에서 사용되는 전략이 이 환경에서는 덜 중요하다는 것을 시사합니다. 데이터는 https://github.com/Leolty/Spider2-AIFunc 에서 확인할 수 있습니다.

Original Abstract

Major cloud data platforms now expose large language model capabilities as native SQL functions, enabling analysts to perform classification, filtering, sentiment analysis, extraction, similarity search, and aggregation within ordinary SQL queries. Yet existing text-to-SQL benchmarks evaluate only conventional SQL and provide no signal on whether models can generate such AI-native SQL. We introduce Spider 2.0-AIFunc, a benchmark of 465 verified instances across 125 real-world databases covering six types of AI functions on the Snowflake platform. Starting from an existing enterprise text-to-SQL benchmark, we construct Spider 2.0-AIFunc through an agent-based pipeline that rewrites source tasks into AI-native form, simultaneously transforming target queries and refining natural language instructions to make the intended AI-native solution explicit and reduce ambiguity. All instances pass a multi-round repeated execution protocol across temporally separated windows to confirm result stability before release. Evaluating ten state-of-the-art language models, we find that the strongest proprietary models reach 67-70% execution accuracy while the best open-source model achieves 58.1%, a gap driven primarily by errors in predicate specification, schema grounding, and AI function parameterization. Agent frameworks designed for traditional text-to-SQL challenges, such as schema retrieval and relevant table selection, do not transfer effectively to AI-native SQL: a minimal agent setup consistently matches or outperforms more elaborate alternatives, suggesting that the strategies these frameworks employ are less critical in this setting. Data are available at https://github.com/Leolty/Spider2-AIFunc .

0 Citations
0 Influential
29.5 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!