PapersWithELO
← ICLR 2024 leaderboard

AN ENTROPY PERSPECTIVE IN KNOWLEDGE DISTILLATION

Jia Guo, Minghao Chen, Yilun Zhao, Boyuan Pan, Yao Hu, Heda Wang, Chen Zhu, Xiaofei He, Deng Cai

self/semi-supervised learningKnowledge Distillation
5.80100
Fused
band ≈ ±15 pct pts (from σ = 0.30)
7.80100
Mimo
band ≈ ±22 pct pts (from σ = 0.44)
6.30100
DeepSeek
band ≈ ±20 pct pts (from σ = 0.40)

OpenReview ground truth

Rejected

Abstract

Knowledge distillation is a widely studied technique for transferring knowledge from a large teacher model to a smaller student model, with the aim of maintaining high performance while reducing computational complexity. However, the performance of the student model often suffers when the teacher model is overly large. We observe significant differences in the ability of teacher and student models to minimize losses, with student models exhibiting higher entropy. This underscores the inherent difficulty in transferring knowledge from the more complex teacher model to the simpler student model. Through theoretical analysis, we propose a straightforward intermediate alignment module to narrow the entropy gap between the student and the teacher, thus enhancing the student performance. Compared with vanilla distillation, the proposed method has the potential to improve the performance of the student model when the teacher model is significantly large, paving the way for more efficient and powerful model learning techniques in the field of knowledge distillation.

Author context

Most prolific author: 5 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Battle history — 34 comparisons

Ranked above opponent in 37% of matchups.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 34)