PapersWithELO
← ICLR 2024 leaderboard

Making Pre-trained Language Models Great on Tabular Prediction

Jiahuan Yan, Bo Zheng, Hongxia Xu, Yiheng Zhu, Danny Chen, Jimeng Sun, Jian Wu, Jintai Chen

transfer & meta learninglanguage modelsclassification and regressionmodel pre-trainingtabular data
72.00100
Fused
band ≈ ±14 pct pts (from σ = 0.28)
67.20100
Mimo
band ≈ ±20 pct pts (from σ = 0.40)
75.60100
DeepSeek
band ≈ ±20 pct pts (from σ = 0.40)

OpenReview ground truth

Accepted

TL;DR — A language model adaption approach for precise tabular data classification and regression.

Abstract

The transferability of deep neural networks (DNNs) has made significant progress in image and language processing. However, due to the heterogeneity among tables, such DNN bonus is still far from being well exploited on tabular data prediction (e.g., regression or classification tasks). Condensing knowledge from diverse domains, language models (LMs) possess the capability to comprehend feature names from various tables, potentially serving as versatile learners in transferring knowledge across distinct tables and diverse prediction tasks, but their discrete text representation space is inherently incompatible with numerical feature values in tables. In this paper, we present TP-BERTa, a specifically pre-trained LM for tabular data prediction. Concretely, a novel relative magnitude tokenization converts scalar numerical feature values to finely discrete, high-dimensional tokens, and an intra-feature attention approach integrates feature values with the corresponding feature names. Comprehensive experiments demonstrate that our pre-trained TP-BERTa leads the performance among tabular DNNs and is competitive with Gradient Boosted Decision Tree models in typical tabular data regime.

Author context

Most prolific author: 14 submissions (credibility 0.57).

Delta if applied: -0.2 percentile

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Battle history — 38 comparisons

Ranked above opponent in 54% of matchups.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 38)