PapersWithELO
← ICLR 2024 leaderboard

ReMasker: Imputing Tabular Data with Masked Autoencoding

Tianyu Du, Luca Melis, Ting Wang

self/semi-supervised learningMasked modelingTabular data imputation
54.70100
Fused
band ≈ ±15 pct pts (from σ = 0.29)
57.30100
Mimo
band ≈ ±20 pct pts (from σ = 0.41)
51.30100
DeepSeek
band ≈ ±21 pct pts (from σ = 0.42)

OpenReview ground truth

Accepted

TL;DR — A simple yet effective imputation method

Abstract

We present ReMasker, a new method of imputing missing values in tabular data by extending the masked autoencoding framework. Compared with prior work, ReMasker is extremely simple -- besides the missing values (i.e., naturally masked), we randomly "re-mask" another set of values, optimize the autoencoder by reconstructing this re-masked set, and apply the trained model to predict the missing values; and yet highly effective -- with extensive evaluation on benchmark datasets, we show that ReMasker performs on par with or outperforms state-of-the-art methods in terms of both imputation fidelity and utility under various missingness settings, while its performance advantage often increases with the ratio of missing data. We further explore theoretical justification for its effectiveness, showing that ReMasker tends to learn missingness-invariant representations of tabular data. Our findings indicate that masked modeling represents a promising direction for further research on tabular data imputation. The code is publicly available.

Author context

Most prolific author: 5 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 34)