Detecting Shortcuts using Mutual Information
Mohammed Adnan, Yani Ioannou, Kenyon Tsai, Angus Galloway, Hamid Tizhoosh, Rahul G Krishnan, Graham W. Taylor
OpenReview ground truth
TL;DR — Proposed a mutual-information based method to detect shortcuts/spurious correlations.
Abstract
The failure of deep neural networks to generalize to out-of-distribution (OOD) data is a well-known problem that raises concerns about the deployment of trained networks in safety-critical domains such as healthcare and autonomous vehicles. We study a particular kind of distribution shift — shortcuts or spurious correlations in the training data. These correlations are not present in real-world test data, so there is a performance drop due to distribution shift, also referred to as shortcut learning. Shortcut learning is often only exposed when models are evaluated in carefully controlled experimental settings, posing a serious dilemma for AI practitioners to properly assess the effectiveness of a trained model for real-world applications. In this work, we try to understand shortcut learning using information-theoretic tools and propose to use the mutual information (MI) between the learned representation and the input space as a domain-agnostic metric for detecting shortcuts in the training datasets. For studying the training dynamics of shortcut learning, we develop a Neural Tangent Kernel (NTK) based framework, which can be used to detect shortcuts and spurious correlations in the training data without requiring class labels of the test data. We empirically demonstrate on multiple datasets, such as MNIST, CelebA, NICO, Waterbirds, and BenchMD, that MI can effectively detect shortcuts. We benchmark against multiple OOD detection baselines to show that OOD detectors cannot detect shortcuts, and our method can be used in complementary with OOD detectors to identify all types of distribution shifts in the datasets, including shortcuts.
Author context
Most prolific author: 3 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 40 comparisons
Ranked above opponent in 50% of matchups.
- ▼ lost to Vibroacoustic Frequency Response Predictio… ×6
- ▼ lost to Improving Private Training via In-distribu… ×6
- ▲ beat FENNs: A Resource-Efficient, Adaptive, Pri… ×4
- ▲ beat Culture in Artificial Intelligence: A Lite… ×4
- ▼ lost to ConjNorm: Tractable Density Estimation for… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 40)