2606.09323v1 Jun 08, 2026 cs.AI

TRL-Bench: 표 데이터 인코더의 패러다임 간, 표현 수준 평가 표준화

TRL-Bench: Standardizing Cross-Paradigm Representation-Level Evaluation of Tabular Encoders

Haoran Xu
Haoran Xu
Citations: 14
h-index: 2
Jinyang Li
Jinyang Li
Citations: 39
h-index: 3
M. T. Ozsu
M. T. Ozsu
Citations: 0
h-index: 0
Xinjian Zhao
Xinjian Zhao
Citations: 4
h-index: 1
Wei Pang
Wei Pang
Citations: 9
h-index: 2
Xiangru Jian
Xiangru Jian
Citations: 4
h-index: 2
Hehan Li
Hehan Li
Citations: 1
h-index: 1
Alex Xue
Alex Xue
Citations: 0
h-index: 0
Zhengyuan Dong
Zhengyuan Dong
University of Waterloo
Citations: 11
h-index: 2
Chao Zhang
Chao Zhang
Citations: 23
h-index: 2
Reynold Cheng
Reynold Cheng
Citations: 485
h-index: 6
Tianshu Yu
Tianshu Yu
Citations: 24
h-index: 4
Zhixuan Yu
Zhixuan Yu
Citations: 41
h-index: 2

표 데이터 인코더는 일반적으로 특정 작업에 최적화된 통합 파이프라인 내에서 평가되므로, 서로 다른 학습 패러다임을 가진 모델들을 유사한 표 데이터를 기반으로 비교하기 어렵습니다. 본 논문에서는 다단계의 표 데이터 표현 학습(TRL) 벤치마크인 TRL-Bench를 소개합니다. TRL-Bench는 다양한 패러다임 간의 표현 수준 평가를 표준화하며, 각 인코더는 지원하는 인터페이스를 통해 행, 열 또는 전체 테이블 임베딩을 출력하고, 공유되는 경량 헤드가 이를 활용하여 세 가지 모듈(TRL-CTbench: 열/테이블, TRL-Rbench: 행, TRL-DLTE: 모든 단계를 포괄하는 복합 데이터 레이크 테이블 풍부화)에서 평가합니다. 이러한 표준화된 환경을 지원하기 위해, 엄선된 벤치마크 자료와 작업 재정의를 제공하며, 여기에는 123개의 검증된 목표를 가진 50개의 OpenML 테이블, 16개의 행 쌍 연결(row-pair linkage) 문제, 그리고 1,379개의 부모 테이블에서 파생된 47,772개의 테이블로 구성된 DLTE 데이터 레이크가 포함됩니다. 20개의 모델과 16개의 작업을 통해 TRL-Bench는 다운스트림 조건이 표준화되면 인코더의 품질은 특정 능력에 따라 결정되며 단일 순위표로 표현될 수 없음을 보여줍니다. TRL-CTbench에서는 일반적인 텍스트 인코더가 표면 텍스트 신호가 강한 작업에서 우수한 성능을 보이는 반면, 표 데이터 전문 모델은 사전 학습 목표가 해당 작업과 일치하는 경우 더 좋은 결과를 얻습니다. TRL-Rbench에서는 테이블 내 예측과 테이블 간 연결이 서로 다른 학습 방식을 선호하며, 원자 단위 연결 성능은 DLTE 파이프라인의 행 매칭 단계와 강한 상관관계를 보입니다. TRL-DLTE에서는 가장 강력한 파이프라인은 특정 능력에 맞는 전문 모델을 결합하는 것이며, 최상의 전체 품질은 개별 단계의 순위뿐만 아니라 비선형적인 복합적 적합성에 달려 있습니다. TRL-Bench는 표준화된 다운스트림 조건 하에서 추출된 표 데이터 표현의 재사용 가능한 신호를 측정하기 위한 공통 프로토콜을 제공합니다. 코드 및 데이터: https://github.com/LOGO-CUHKSZ/TRL-Bench

Original Abstract

Tabular encoders are usually evaluated inside task-specific end-to-end pipelines, so models from different training paradigms are difficult to compare directly even when they operate on similar tabular signals. We introduce TRL-Bench, a multi-granular tabular representation learning (TRL) benchmark that standardizes cross-paradigm representation-level evaluation: each encoder exports row-, column-, or table embeddings through its supported wrapper, and shared lightweight heads probe them across three suites: TRL-CTbench (column/table), TRL-Rbench (row), and TRL-DLTE (compositional Data-Lake Table Enrichment spanning all three granularities). To support this standardized setting, we release curated benchmark assets and task reformulations, including 50 OpenML tables with 123 verified targets, 16 row-pair linkage rewrites, and a 47,772-table DLTE lake derived from 1,379 parent tables. Across 20 models and 16 tasks, TRL-Bench shows that once downstream conditions are standardized, encoder quality is capability-specific rather than captured by a single leaderboard. In TRL-CTbench, generic text encoders often lead on tasks with strong surface-text signal, while tabular specialists win where their pretraining objective aligns with the task. In TRL-Rbench, within-table prediction and cross-table linkage favor different training regimes, with atomic linkage performance correlating strongly with the row-matching stage of DLTE pipelines. In TRL-DLTE, the strongest pipelines combine capability-matched specialists rather than reuse a single encoder, and top end-to-end quality depends on non-additive compositional fit rather than per-stage marginal rank alone. TRL-Bench provides a common protocol for measuring reusable signal in exported tabular representations under shared downstream conditions. Code and data: https://github.com/LOGO-CUHKSZ/TRL-Bench

1 Citations
0 Influential
28.493061443341 Altmetric
6.9 Score
Original PDF
2

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!