AN ENTROPY PERSPECTIVE IN KNOWLEDGE DISTILLATION
Jia Guo, Minghao Chen, Yilun Zhao, Boyuan Pan, Yao Hu, Heda Wang, Chen Zhu, Xiaofei He, Deng Cai
OpenReview ground truth
Abstract
Knowledge distillation is a widely studied technique for transferring knowledge from a large teacher model to a smaller student model, with the aim of maintaining high performance while reducing computational complexity. However, the performance of the student model often suffers when the teacher model is overly large. We observe significant differences in the ability of teacher and student models to minimize losses, with student models exhibiting higher entropy. This underscores the inherent difficulty in transferring knowledge from the more complex teacher model to the simpler student model. Through theoretical analysis, we propose a straightforward intermediate alignment module to narrow the entropy gap between the student and the teacher, thus enhancing the student performance. Compared with vanilla distillation, the proposed method has the potential to improve the performance of the student model when the teacher model is significantly large, paving the way for more efficient and powerful model learning techniques in the field of knowledge distillation.
Author context
Most prolific author: 5 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 34 comparisons
Ranked above opponent in 37% of matchups.
- ▲ beat OWL: A Large Language Model for IT Operati… ×6
- ▲ beat DAG-based Generative Regression ×6
- ▲ beat Rethinking the Buyer’s Inspection Paradox … ×6
- ▲ beat Impact of Molecular Representations on Dee… ×6
- ▼ lost to Learning Hierarchical Image Segmentation F… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 34)