2604.06814v1 Apr 08, 2026 cs.LG

OmniTabBench: 테이블 데이터에 대한 GBDT, 신경망 및 기반 모델의 경험적 한계를 대규모로 분석

OmniTabBench: Mapping the Empirical Frontiers of GBDTs, Neural Networks, and Foundation Models for Tabular Data at Scale

Di Jiang
Di Jiang
Citations: 63
h-index: 1
Ruoqi Cao
Ruoqi Cao
Citations: 12
h-index: 2
Zhiyuan Dang
Zhiyuan Dang
Citations: 384
h-index: 8
Li Huang
Li Huang
Citations: 1
h-index: 1
Qingsong Zhang
Qingsong Zhang
Citations: 188
h-index: 6
Shihao Piao
Shihao Piao
Citations: 1
h-index: 1
Shenggao Zhu
Shenggao Zhu
Citations: 1,079
h-index: 16
Zhouchen Lin
Zhouchen Lin
Citations: 12
h-index: 2
Qi Tian
Qi Tian
Citations: 762
h-index: 18
Zhiyu Wang
Zhiyu Wang
Citations: 1
h-index: 1
Jianlong Chang
Jianlong Chang
Citations: 58
h-index: 2

전통적인 트리 기반 앙상블 방법이 오랫동안 테이블 데이터 관련 작업에서 우위를 점해 왔지만, 딥 신경망과 새로운 기반 모델이 이러한 지배력을 위협하고 있습니다. 그러나 보편적으로 우수한 패러다임에 대한 합의는 아직 없습니다. 기존 벤치마크는 일반적으로 100개 미만의 데이터 세트를 포함하며, 이는 평가의 충분성 및 잠재적인 선택 편향에 대한 우려를 야기합니다. 이러한 제한 사항을 해결하기 위해, 우리는 현재까지 가장 큰 테이블 데이터 벤치마크인 OmniTabBench를 소개합니다. OmniTabBench는 다양한 작업 범위를 포괄하는 3030개의 데이터 세트로 구성되어 있으며, 대규모 언어 모델을 사용하여 다양한 소스에서 수집하고 산업 분야별로 분류되었습니다. 우리는 OmniTabBench에서 모든 모델 패밀리의 최첨단 모델에 대한 전례 없는 대규모 경험적 평가를 수행하여, 특정 모델이 압도적으로 우세하지 않음을 확인했습니다. 또한, 데이터 세트 크기, 특징 유형, 특징 및 목표 변수의 왜도/첨도와 같은 개별 속성을 분석하는 분리된 메타 특징 분석을 통해 특정 모델 범주에 유리한 조건을 규명했습니다. 이는 기존의 복합 지표 연구보다 더 명확하고 실용적인 지침을 제공합니다.

Original Abstract

While traditional tree-based ensemble methods have long dominated tabular tasks, deep neural networks and emerging foundation models have challenged this primacy, yet no consensus exists on a universally superior paradigm. Existing benchmarks typically contain fewer than 100 datasets, raising concerns about evaluation sufficiency and potential selection biases. To address these limitations, we introduce OmniTabBench, the largest tabular benchmark to date, comprising 3030 datasets spanning diverse tasks that are comprehensively collected from diverse sources and categorized by industry using large language models. We conduct an unprecedented large-scale empirical evaluation of state-of-the-art models from all model families on OmniTabBench, confirming the absence of a dominant winner. Furthermore, through a decoupled metafeature analysis, which examines individual properties such as dataset size, feature types, feature and target skewness/kurtosis, we elucidate conditions favoring specific model categories, providing clearer, more actionable guidance than prior compound-metric studies.

2 Citations
0 Influential
9 Altmetric
47.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!