PapersWithELO
← ICLR 2024 leaderboard

Mechanism of clean-priority learning in early stopped neural networks of infinite width

Chaoyue Liu, Amirhesam Abedsoltan, Mikhail Belkin

learning theorylabel noiseearly stoppingclean-priority learning
75.00100
Fused
band ≈ ±15 pct pts (from σ = 0.29)
76.30100
Mimo
band ≈ ±21 pct pts (from σ = 0.42)
70.90100
DeepSeek
band ≈ ±20 pct pts (from σ = 0.40)

OpenReview ground truth

Rejected

TL;DR — We theoretically disclose the underlying mechanism and dynamics that are responsible for the clean-priority learning (and its termination), via analysis of sample-wise gradients of infinitely wide neural networks

Abstract

When random label noise is added to a training dataset, the prediction error of a neural network on a label-noise-free test dataset initially improves during early training but eventually deteriorates, following a U-shaped dependence on training time. This behaviour is believed to be a result of neural networks learning the pattern of clean data first and fitting the noise later, a phenomenon that we refer to as *clean-priority learning*. In this study, we aim to explore the learning dynamics underlying this phenomenon. We demonstrate that, in the early stage of training, the update direction of gradient descent is determined by the clean samples of training data, leaving the noisy samples have minimal to no impact, resulting in a prioritization of clean learning. Moreover, we show both theoretically and experimentally, as the clean-priority learning goes on, the dominance of the gradients of clean samples over those of noisy samples diminishes, and finally results in a termination of the clean-priority learning and fitting of the noisy samples.

Author context

Most prolific author: 5 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 38)