PapersWithELO
← ICLR 2024 leaderboard

Conditional MAE: An Empirical Study of Multiple Masking in Masked Autoencoder

Jie Zhu, Zhihao Yu, Mingyu Ding, Ping Luo, Leye Wang

interpretability & vizmasked autoencodermultiple masking
17.10100
Fused
band ≈ ±14 pct pts (from σ = 0.29)
16.50100
Mimo
band ≈ ±21 pct pts (from σ = 0.41)
11.70100
DeepSeek
band ≈ ±20 pct pts (from σ = 0.40)

OpenReview ground truth

Rejected

Abstract

This work aims to study the subtle yet often overlooked element of masked autoencoder (MAE): masking. While masking plays a critical role in the performance of MAE, most current research employs fixed masking strategies directly on the input image. We introduce a masked autoencoder framework with multiple masking stages, termed Conditional MAE, where subsequent maskings are conditioned on previous unmasked representations, enabling a more flexible masking process in masked image modeling. By doing so, our study sheds light on how multiple masking affects the optimization in training and performance of pretrained models, e.g., introducing more locality to models, and summarizes several takeaways from our findings. Finally, we empirically evaluate the performance of our best-performing model (Conditional-MAE) with that of MAE in three folds including transfer learning, robustness, and scalability, demonstrating the effectiveness of our multiple masking strategy. We hope our findings will inspire further research in the field and code will be made available.

Author context

Most prolific author: 19 submissions (credibility 0.23).

Delta if applied: -1.1 percentile

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 34)