$\nu$-ensembles: Improving deep ensemble calibration in the small data regime
Konstantinos Pitas, Julyan Arbel
OpenReview ground truth
TL;DR — We use unlabeled data to improve deep ensemble diversity and calibration for small to medium-sized training sets.
Abstract
We present a method to improve the calibration of deep ensembles in the small data regime in the presence of unlabeled data. Our approach, which we name $\nu$-ensembles, is extremely easy to implement: given an unlabeled set, for each unlabeled data point, we simply fit a different randomly selected label with each ensemble member. We provide a theoretical analysis based on a PAC-Bayes bound which guarantees that for such a labeling we obtain low negative log-likelihood and high ensemble diversity on testing samples. Empirically, through detailed experiments, we find that for low to moderately-sized training sets, $\nu$-ensembles are more diverse and provide better calibration than standard ensembles, sometimes significantly.
Author context
Most prolific author: 1 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 34 comparisons
Ranked above opponent in 48% of matchups.
- ▲ beat Maximum Entropy On-Policy Actor-Critic via… ×6
- ▲ beat Towards Reliable Evaluation and Fast Train… ×4
- ▲ beat Evaluating the Evaluators: Are Current Few… ×4
- ▼ lost to MT-Ranker: Reference-free machine translat… ×4
- ▼ lost to Calibration Attack: A Framework For Advers… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 34)