← ICLR 2024 leaderboard

Cosine Similarity Knowledge Distillation for Individual Class Information Transfer

Gyeongdo Ham, Seonghak KIM, Suin Lee, Jae-Hyeok Lee, Daeshik Kim

representation learningDeep learningimage classificationknowledge distillationcosine similaritycosine distance
11.90100
Fused
band ≈ ±13 pct pts (from σ = 0.27)
14.10100
Mimo
band ≈ ±20 pct pts (from σ = 0.39)
11.40100
DeepSeek
band ≈ ±18 pct pts (from σ = 0.37)

OpenReview ground truth

Rejected

TL;DR — We introduce a novel loss function that enables the student model to dynamically acquire the teacher knowledge.

Abstract

Previous logits-based Knowledge Distillation (KD) have utilized predictions about multiple categories within each sample (i.e., class predictions) and have employed Kullback-Leibler (KL) divergence to reduce the discrepancy between the student’s and teacher’s predictions. Despite the proliferation of KD techniques, the student model continues to fall short of achieving a similar level as teachers. In response, we introduce a novel and effective KD method capable of achieving results on par with or superior to the teacher model’s performance. We utilize teacher and student predictions about multiple samples for each category (i.e., batch predictions) and apply cosine similarity, a commonly used technique in Natural Language Processing (NLP) for measuring the resemblance between text embeddings. This metric's inherent scale-invariance property, which relies solely on vector direction and not magnitude, allows the student to dynamically learn from the teacher's knowledge, rather than being bound by a fixed distribution of the teacher's knowledge. Furthermore, we propose a method called cosine similarity weighted temperature (CSWT) to improve the performance. CSWT reduces the temperature scaling in KD when the cosine similarity between the student and teacher models is high, and conversely, it increases the temperature scaling when the cosine similarity is low. This adjustment optimizes the transfer of information from the teacher to the student model. Extensive experimental results show that our proposed method serves as a viable alternative to existing methods. We anticipate that this approach will offer valuable insights for future research on model compression.

Author context

Most prolific author: 1 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 42)