PapersWithELO
← ICLR 2024 leaderboard

An improved analysis of per-sample and per-update clipping in federated learning

Bo Li, Xiaowen Jiang, Mikkel N. Schmidt, Tommy Sonne Alstrøm, Sebastian U Stich

optimizationclippingfederated learningdecentralized learningdistributed optimization
64.20100
Fused
band ≈ ±14 pct pts (from σ = 0.28)
62.10100
Mimo
band ≈ ±21 pct pts (from σ = 0.42)
69.00100
DeepSeek
band ≈ ±19 pct pts (from σ = 0.37)

OpenReview ground truth

Accepted

TL;DR — We rigorously and precisely analyze the impact of per-sample and per-update clipping on the convergence of FedAvg

Abstract

Gradient clipping is key mechanism that is essential to differentially private training techniques in Federated learning. Two popular strategies are per-sample clipping, which clips the mini-batch gradient, and per-update clipping, which clips each user's model update. However, there has not been a thorough theoretical analysis of these two clipping methods. In this work, we rigorously analyze the impact of these two clipping techniques on the convergence of a popular federated learning algorithm FedAvg under standard stochastic noise and gradient dissimilarity assumptions. We provide a convergence guarantee given any arbitrary clipping threshold. Specifically, we show that per-sample clipping is guaranteed to converge to the neighborhood of the stationary point, with the size dependent on the stochastic noise, gradient dissimilarity, and clipping threshold. In contrast, the convergence to the stationary point can be guaranteed with a sufficiently small stepsize in per-update clipping at the cost of more communication rounds. We further provide insights into understanding the impact of the improved convergence analysis in the differentially private setting.

Author context

Most prolific author: 5 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Battle history — 38 comparisons

Ranked above opponent in 54% of matchups.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 38)