PapersWithELO
← ICLR 2024 leaderboard

Calibration Bottleneck: What Makes Neural Networks less Calibratable?

Deng-Bao Wang, Min-Ling Zhang

self/semi-supervised learningUncertainty CalibrationPost-hoc Calibration
47.30100
Fused
band ≈ ±14 pct pts (from σ = 0.28)
45.00100
Mimo
band ≈ ±19 pct pts (from σ = 0.37)
50.60100
DeepSeek
band ≈ ±20 pct pts (from σ = 0.41)

OpenReview ground truth

Rejected

Abstract

While modern deep neural networks have achieved remarkable success, they have exhibited a notable deficiency in reliably estimating uncertainty. Many existing studies address the uncertainty calibration problem by incorporating regularization techniques to penalize the overconfident outputs during training. In this study, we shift the focus from the miscalibration encountered in the training phase to an investigation of the concept of calibratability, assessing how amenable a model is to be recalibrated in post-training phase. We find that the use of regularization techniques might compromise calibratability, subsequently leading to a decline in final calibration performance after recalibration. To identify the underlying causes leading to poor calibratability, we delve into the calibration of intermediate features across neural networks’ hidden layers. Our study demonstrates that the overtraining of the top layers in neural networks poses a significant obstacle to calibration, while these layers typically offer minimal improvement to the discriminability of features. Based on this observation, we introduce a weak classifier hypothesis: Given a weak classification head, the bottom layers of a neural network can be learned better for producing calibratable features. Consequently, we propose a progressively layer-peeled training (PLT) method to exploit this hypothesis, thereby enhancing model calibratability. Comprehensive experiments show the effectiveness of our method, which improves model calibration and also yields competitive predictive performance.

Author context

Most prolific author: 2 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 38)