Personalization Mitigates the Perils of Local SGD for Heterogeneous Distributed Learning
Kumar Kshitij Patel, Nidham Gazagnadou, Lingxiao Wang, Lingjuan Lyu
OpenReview ground truth
TL;DR — Improved rates for personalized local SGD dominating all existing baselines without data-heterogeneity assumptions.
Abstract
This paper investigates a personalized version of Local Stochastic Gradient Descent (Local SGD). We establish improved convergence guarantees for this personalized approach, eliminating the need for extra assumptions about data or gradient heterogeneity. Our theoretical analysis reveals that personalized Local SGD outperforms both pure local training and federated learning algorithms that produce a consensus model for all devices. This performance gain is primarily due to over-parameterization, which allows for reducing the consensus error between clients with more communication—something that is not observed in non-personalized approaches. We illustrate our observations using experiments on synthetic convex and smooth objectives.
Author context
Most prolific author: 11 submissions (credibility 0.92).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 36 comparisons
Ranked above opponent in 57% of matchups.
- ▼ lost to Optimal Sketching for Residual Error Estim… ×4
- ▼ lost to Diffusion Denoising as a Certified Defense… ×4
- ▼ lost to Revisiting High-Resolution ODEs for Faster… ×4
- ▼ lost to A Theoretical Study of the Jacobian Matrix… ×4
- ▼ lost to PREDICTING ACCURATE LAGRANGIAN MULTIPLIERS… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 36)