PapersWithELO
← ICLR 2024 leaderboard

$\nu$-ensembles: Improving deep ensemble calibration in the small data regime

Konstantinos Pitas, Julyan Arbel

fairness, safety & privacydeep ensemblescalibrationuncertaintydiversityPAC-Bayes
46.60100
Fused
band ≈ ±14 pct pts (from σ = 0.29)
40.70100
Mimo
band ≈ ±21 pct pts (from σ = 0.42)
65.80100
DeepSeek
band ≈ ±20 pct pts (from σ = 0.40)

OpenReview ground truth

Rejected

TL;DR — We use unlabeled data to improve deep ensemble diversity and calibration for small to medium-sized training sets.

Abstract

We present a method to improve the calibration of deep ensembles in the small data regime in the presence of unlabeled data. Our approach, which we name $\nu$-ensembles, is extremely easy to implement: given an unlabeled set, for each unlabeled data point, we simply fit a different randomly selected label with each ensemble member. We provide a theoretical analysis based on a PAC-Bayes bound which guarantees that for such a labeling we obtain low negative log-likelihood and high ensemble diversity on testing samples. Empirically, through detailed experiments, we find that for low to moderately-sized training sets, $\nu$-ensembles are more diverse and provide better calibration than standard ensembles, sometimes significantly.

Author context

Most prolific author: 1 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 34)