Conditional MAE: An Empirical Study of Multiple Masking in Masked Autoencoder
Jie Zhu, Zhihao Yu, Mingyu Ding, Ping Luo, Leye Wang
OpenReview ground truth
Abstract
This work aims to study the subtle yet often overlooked element of masked autoencoder (MAE): masking. While masking plays a critical role in the performance of MAE, most current research employs fixed masking strategies directly on the input image. We introduce a masked autoencoder framework with multiple masking stages, termed Conditional MAE, where subsequent maskings are conditioned on previous unmasked representations, enabling a more flexible masking process in masked image modeling. By doing so, our study sheds light on how multiple masking affects the optimization in training and performance of pretrained models, e.g., introducing more locality to models, and summarizes several takeaways from our findings. Finally, we empirically evaluate the performance of our best-performing model (Conditional-MAE) with that of MAE in three folds including transfer learning, robustness, and scalability, demonstrating the effectiveness of our multiple masking strategy. We hope our findings will inspire further research in the field and code will be made available.
Author context
Most prolific author: 19 submissions (credibility 0.23).
Delta if applied: -1.1 percentile
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 34 comparisons
Ranked above opponent in 44% of matchups.
- ▲ beat Rephrase, Augment, Reason: Visual Groundin… ×6
- ▼ lost to Unifying Feature and Cost Aggregation with… ×4
- ▼ lost to A Backdoor-based Explainable AI Benchmark … ×4
- ▼ lost to On the Joint Interaction of Models, Data, … ×4
- ▲ beat Computing high-dimensional optimal transpo… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 34)