DataSpace: 다양한 작업 환경에서의 검증 가능한 분석을 위한 데이터 에이전트 성능 평가
DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces
데이터 에이전트는 조직 내 작업 공간에서 자연어 기반의 분석을 가능하게 하며, 관련 증거는 데이터베이스, 정형 파일, 긴 문서, 멀티미디어 등 다양한 형태로 분산되어 있을 수 있습니다. 기존의 벤치마크들은 주로 구조화된 질의, 검색 또는 개방형 분석에 초점을 맞추고 있어, 이질적인 증거 탐색, 완전한 테이블 형태의 결과 생성 및 결정론적인 평가를 통합적으로 다루지 못합니다. 본 논문에서는 데이터 에이전트가 작업별로 분산된 다양한 환경에서 검증 가능한 테이블 형태의 결과를 생성하도록 설계된 벤치마크인 DataSpace를 소개합니다. DataSpace는 410개의 다국어 작업을 포함하며, CSV, JSON, SQLite, Markdown, PDF 및 비디오 형식으로 총 7,439개의 데이터 세트를 포함하고 있으며, 전체 용량은 15.01 GB입니다. 또한, DataSpace는 KDD Cup 2026의 복잡한 데이터 분석을 위한 데이터 에이전트 대회 공식 평가 벤치마크로 사용되었습니다. 각 에이전트는 질문과 작업 환경만을 입력으로 받아 요청된 완전한 테이블 형태의 결과를 반환합니다. DataSpace는 DataSpace-Builder라는 실행 기반 프레임워크를 사용하여 구축되었으며, 이 프레임워크는 다국어 변환, 제약 조건 기반 관계형 샘플링, 모달리티 라우팅 및 아티팩트 렌더링 기능을 포함하며, 11명의 분야 전문가에 의해 검토 및 수정되었습니다. 결정론적 평가 도구는 헤더 불변 컬럼 정렬, 타입 및 정밀도 인지 정규화, 그리고 순서 기반 행 비교를 수행합니다. 최근 출시된 6개의 멀티모달 모델과 널리 사용되는 5가지 에이전트 프레임워크를 사용하여 성능을 평가한 결과, 가장 높은 정확도는 66.34%였습니다. 또한, 프레임워크 선택에 따라 정확도가 최대 15.36 포인트 차이를 보였습니다. 멀티모달 증거 통합 및 조인은 모든 6개의 모델에서 일관적으로 정확도를 감소시키는 것으로 나타났습니다. 이러한 결과는 DataSpace가 아직 발전 가능성이 높으며, 데이터 에이전트의 신뢰성을 향상시키기 위한 주요 과제를 제시하고 있음을 보여줍니다.
Data agents enable natural-language analytics over organizational workspaces, where relevant evidence may be scattered across databases, structured files, long documents, and multimedia. Existing benchmarks largely isolate structured querying, retrieval, or open-ended analysis, leaving heterogeneous evidence discovery, complete tabular outputs, and deterministic evaluation insufficiently unified. We introduce DataSpace, a benchmark in which data agents produce verifiable tabular results from task-local heterogeneous workspaces. It contains 410 cross-language tasks and 7,439 artifacts totaling 15.01 GB across CSV, JSON, SQLite, Markdown, PDF, and video. DataSpace also served as the official evaluation benchmark for the KDD Cup 2026 Data Agents for Complex Data Analysis competition. Each agent receives only a question and workspace and returns the complete requested tabular result. We construct DataSpace with DataSpace-Builder, an execution-grounded framework comprising cross-language transformation, constraint-aware relational sampling, modality routing and artifact rendering, and human review and task repair by 11 domain experts. A deterministic evaluator performs header-invariant column alignment, type- and precision-aware normalization, and order-aware row comparison. Across six recently released frontier multimodal models and five widely used agent harnesses, the best accuracy reaches 66.34%, while harness choice creates a 15.36-point spread with the backbone fixed. Multimodal evidence integration and joins consistently reduce accuracy across all six backbones. These results show that DataSpace remains unsaturated and identify key challenges for improving data-agent reliability.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.