PapersWithELO
← ICLR 2024 leaderboard

Image Compression Is an Effective Objective for Visual Representation Learning

Yunjie Tian, Lingxi Xie

self/semi-supervised learningVisual Pre-trainingSelf-supervised LearningImage CompressionKolmogorov Complexity
38.70100
Fused
band ≈ ±14 pct pts (from σ = 0.29)
34.30100
Mimo
band ≈ ±20 pct pts (from σ = 0.40)
49.30100
DeepSeek
band ≈ ±21 pct pts (from σ = 0.41)

OpenReview ground truth

Rejected

TL;DR — A novel visual pre-training methodology based on image data compression, which is degradation-free and applies to various vision transformers

Abstract

Self-supervised pre-training is an effective method for initializing the weights of vision transformers. In this paper, we advocate for a novel learning objective that trains the target model to use a minimal number of tokens to reconstruct images. Compared to the existing approaches including contrastive learning (CL) and masked image modeling (MIM), our formulation not only offers a new perspective of visual pre-training from the information theory, but also alleviates the degradation dilemma which may lead to instability. The idea is implemented using Semantic Merging and Reconstruction (SMR). SMR feeds the entire image (without any degradation) into the target model, gradually reduces the number of tokens throughout the encoder, and requires the decoder to maximally recover the original image in the semantic space using the remaining tokens. We establish SMR upon the vanilla ViT and two of its variants. Under the standard evaluation protocol, SMR shows favorable performance in visual pre-training and various downstream tasks. Additionally, SMR enjoys reduced pre-training time and memory consumption and thus is scalable to pre-train very large vision models. Code is submitted as supplementary material and will be open-sourced.

Author context

Most prolific author: 2 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 34)