Mechanism of clean-priority learning in early stopped neural networks of infinite width
Chaoyue Liu, Amirhesam Abedsoltan, Mikhail Belkin
OpenReview ground truth
TL;DR — We theoretically disclose the underlying mechanism and dynamics that are responsible for the clean-priority learning (and its termination), via analysis of sample-wise gradients of infinitely wide neural networks
Abstract
When random label noise is added to a training dataset, the prediction error of a neural network on a label-noise-free test dataset initially improves during early training but eventually deteriorates, following a U-shaped dependence on training time. This behaviour is believed to be a result of neural networks learning the pattern of clean data first and fitting the noise later, a phenomenon that we refer to as *clean-priority learning*. In this study, we aim to explore the learning dynamics underlying this phenomenon. We demonstrate that, in the early stage of training, the update direction of gradient descent is determined by the clean samples of training data, leaving the noisy samples have minimal to no impact, resulting in a prioritization of clean learning. Moreover, we show both theoretically and experimentally, as the clean-priority learning goes on, the dominance of the gradients of clean samples over those of noisy samples diminishes, and finally results in a termination of the clean-priority learning and fitting of the noisy samples.
Author context
Most prolific author: 5 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 38 comparisons
Ranked above opponent in 60% of matchups.
- ▲ beat Metanetwork: A novel approach to interpret… ×4
- ▲ beat Understanding Deep Neural Networks as Dyna… ×4
- ▼ lost to On the Provable Advantage of Unsupervised … ×4
- ▲ beat Generalized Convergence Analysis of Tsetli… ×4
- ▼ lost to Learning to Reject Meets Long-tail Learnin… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 38)