PapersWithELO
← ICLR 2024 leaderboard

Personalization Mitigates the Perils of Local SGD for Heterogeneous Distributed Learning

Kumar Kshitij Patel, Nidham Gazagnadou, Lingxiao Wang, Lingjuan Lyu

optimizationFederated LearningOptimizationData HeterogenenityLocal SGDFederated AveragingPersonalizationConvergence AnalysisPrivacy Preserving Machine Learning
69.50100
Fused
band ≈ ±15 pct pts (from σ = 0.30)
73.00100
Mimo
band ≈ ±21 pct pts (from σ = 0.43)
75.90100
DeepSeek
band ≈ ±21 pct pts (from σ = 0.42)

OpenReview ground truth

Rejected

TL;DR — Improved rates for personalized local SGD dominating all existing baselines without data-heterogeneity assumptions.

Abstract

This paper investigates a personalized version of Local Stochastic Gradient Descent (Local SGD). We establish improved convergence guarantees for this personalized approach, eliminating the need for extra assumptions about data or gradient heterogeneity. Our theoretical analysis reveals that personalized Local SGD outperforms both pure local training and federated learning algorithms that produce a consensus model for all devices. This performance gain is primarily due to over-parameterization, which allows for reducing the consensus error between clients with more communication—something that is not observed in non-personalized approaches. We illustrate our observations using experiments on synthetic convex and smooth objectives.

Author context

Most prolific author: 11 submissions (credibility 0.92).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Battle history — 36 comparisons

Ranked above opponent in 57% of matchups.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 36)