An improved analysis of per-sample and per-update clipping in federated learning
Bo Li, Xiaowen Jiang, Mikkel N. Schmidt, Tommy Sonne Alstrøm, Sebastian U Stich
OpenReview ground truth
TL;DR — We rigorously and precisely analyze the impact of per-sample and per-update clipping on the convergence of FedAvg
Abstract
Gradient clipping is key mechanism that is essential to differentially private training techniques in Federated learning. Two popular strategies are per-sample clipping, which clips the mini-batch gradient, and per-update clipping, which clips each user's model update. However, there has not been a thorough theoretical analysis of these two clipping methods. In this work, we rigorously analyze the impact of these two clipping techniques on the convergence of a popular federated learning algorithm FedAvg under standard stochastic noise and gradient dissimilarity assumptions. We provide a convergence guarantee given any arbitrary clipping threshold. Specifically, we show that per-sample clipping is guaranteed to converge to the neighborhood of the stationary point, with the size dependent on the stochastic noise, gradient dissimilarity, and clipping threshold. In contrast, the convergence to the stationary point can be guaranteed with a sufficiently small stepsize in per-update clipping at the cost of more communication rounds. We further provide insights into understanding the impact of the improved convergence analysis in the differentially private setting.
Author context
Most prolific author: 5 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 38 comparisons
Ranked above opponent in 54% of matchups.
- ▼ lost to An Embodied Generalist Agent in 3D World ×6
- ▼ lost to Provably Doubly Accelerated Federated Lear… ×4
- ▼ lost to Federated Learning, Lessons from Generaliz… ×4
- ▼ lost to Rethinking Information-theoretic Generaliz… ×4
- ▲ beat Analyzing Implicit Regularization In Feder… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 38)