Gradient norm as a powerful proxy to out-of-distribution error estimation
RENCHUNZI XIE, Ambroise Odonnat, Vasilii Feofanov, Ievgen Redko, Jianfeng Zhang, Bo An
OpenReview ground truth
TL;DR — We use the norm of classification-layer gradients, backpropagated from the cross entropy loss with only one gradient step over OOD data, to formulate an estimation score correlating with the expected OOD error.
Abstract
Estimating out-of-distribution (OOD) error without access to the ground-truth test labels is a highly challenging, yet extremely important problem in the safe deployment of machine learning algorithms. Current works rely on the information from either the outputs or the extracted features to formulate an estimation score correlating with the expected OOD error. In this paper, we investigate--both empirically and theoretically--how the information provided by the gradients can be predictive of the OOD error. Specifically, we use the norm of classification-layer gradients, backpropagated from the cross-entropy loss with only one gradient step over OOD data. Our key idea is that the model should be adjusted with a higher magnitude of gradients when it does not generalize to the OOD dataset. We provide theoretical insights highlighting the main ingredients of such an approach ensuring its empirical success. Extensive experiments conducted on diverse distribution shifts and model structures demonstrate that our method outperforms state-of-the-art algorithms significantly.
Author context
Most prolific author: 14 submissions (credibility 0.79).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 30 comparisons
Ranked above opponent in 51% of matchups.
- ▲ beat Towards Efficient Trace Estimation for Opt… ×4
- ▼ lost to On Representation Complexity of Model-base… ×4
- ▲ beat BOWLL: A DECEPTIVELY SIMPLE OPEN WORLD LIF… ×4
- ▼ lost to Detecting, Explaining, and Mitigating Memo… ×4
- ▲ beat MacDC: Masking-augmented Collaborative Dom… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 30)