Image Compression Is an Effective Objective for Visual Representation Learning
Yunjie Tian, Lingxi Xie
OpenReview ground truth
TL;DR — A novel visual pre-training methodology based on image data compression, which is degradation-free and applies to various vision transformers
Abstract
Self-supervised pre-training is an effective method for initializing the weights of vision transformers. In this paper, we advocate for a novel learning objective that trains the target model to use a minimal number of tokens to reconstruct images. Compared to the existing approaches including contrastive learning (CL) and masked image modeling (MIM), our formulation not only offers a new perspective of visual pre-training from the information theory, but also alleviates the degradation dilemma which may lead to instability. The idea is implemented using Semantic Merging and Reconstruction (SMR). SMR feeds the entire image (without any degradation) into the target model, gradually reduces the number of tokens throughout the encoder, and requires the decoder to maximally recover the original image in the semantic space using the remaining tokens. We establish SMR upon the vanilla ViT and two of its variants. Under the standard evaluation protocol, SMR shows favorable performance in visual pre-training and various downstream tasks. Additionally, SMR enjoys reduced pre-training time and memory consumption and thus is scalable to pre-train very large vision models. Code is submitted as supplementary material and will be open-sourced.
Author context
Most prolific author: 2 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 34 comparisons
Ranked above opponent in 45% of matchups.
- ▼ lost to Rethinking Self-Supervise Learning: An Ins… ×6
- ▲ beat Safeguarding Data in Multimodal AI: A Diff… ×4
- ▲ beat Large Scene Synthesis Controlled With Deta… ×4
- ▼ lost to Patched Denoising Diffusion Models For Hig… ×4
- ▼ lost to On robust overfitting: adversarial trainin… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 34)