Analyzing Implicit Regularization In Federated Learning
Jinwoo Lim, Sangyoon Yu, Suhyun Kim, Soo-Mook Moon
OpenReview ground truth
Abstract
Backward error analysis is a powerful technique that can check how much the path of the gradient flow is modified under the influence of a finite learning rate. Through this technique, it is also possible to find an implicit regularizer that affects the convergence behavior of an optimizer. With a backward error analysis, this paper seeks a more intuitive but quantitative way to understand the convergence behaviour under various federated learning algorithms. We prove that the implicit regularizer for FedAvg disperses the gradient of each client from the average gradient, increasing the gradient variance. We then theoretically present that the implicit regularizer of FedAvg hampers the convergence if the variance of gradients from clients decreases following the gradient of the cost function. In order to verify our analysis, we run experiments on FedAvg with and without the drifting term and confirm that FedAvg without the drifting term shows higher test accuracies. Our analysis also explains the convergence behavior of variance reduction methods such as SCAFFOLD, FedDyn, and FedSAM to show that the implicit regularizers of those methods have a smaller or zero drifting effect when the learning rate is small. Especially, we provide a possible reason FedSAM can perform better than FedAvg but might not perform as well as other stable variance reduction methods under data heterogeneity.
Author context
Most prolific author: 3 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 38 comparisons
Ranked above opponent in 50% of matchups.
- ▼ lost to Provably Doubly Accelerated Federated Lear… ×4
- ▼ lost to Federated Learning, Lessons from Generaliz… ×4
- ▼ lost to Rethinking Information-theoretic Generaliz… ×4
- ▼ lost to An improved analysis of per-sample and per… ×4
- ▲ beat Benchmarking Large Language Models as AI R… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 38)