PapersWithELO
← ICLR 2024 leaderboard

Retro: Reusing teacher projection head for efficient embedding distillation on Lightweight Models via Self-supervised Learning

Khanh-Binh Nguyen

self/semi-supervised learningSelf-supervised learningknowledge distillationlightweight modelscontrastive learningconsistency learning
41.10100
Fused
band ≈ ±13 pct pts (from σ = 0.26)
44.00100
Mimo
band ≈ ±19 pct pts (from σ = 0.39)
37.00100
DeepSeek
band ≈ ±18 pct pts (from σ = 0.36)

OpenReview ground truth

Rejected

TL;DR — Reusing teacher projection head for efficient lightweight model distillation via SSL

Abstract

Self-supervised learning (SSL) is gaining attention for its ability to learn effective representations with large amounts of unlabeled data. Lightweight models can be distilled from larger self-supervised pre-trained models using contrastive and consistency constraints, but the different sizes of the projection heads make it challenging for students to accurately mimic the teacher's embedding. We propose \textsc{Retro}, which reuses the teacher's projection head for students, and our experimental results demonstrate significant improvements over the state-of-the-art on all lightweight models. For instance, when training EfficientNet-B0 using ResNet-50/101/152 as teachers, our approach improves the linear result on ImageNet to $66.9%$, $69.3%$, and $69.8%$, respectively, with significantly fewer parameters.

Author context

Most prolific author: 1 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 46)