Making Batch Normalization Great in Federated Deep Learning
Jike Zhong, Hong-You Chen, Wei-Lun Chao
OpenReview ground truth
TL;DR — We investigate the cause of BN issue in FL and introduce a straightforward solution that fully recovers BN failure.
Abstract
Batch Normalization (BN) is commonly used in modern deep learning to improve stability and speed up convergence in centralized training. In federated learning (FL) with non-IID decentralized data, previous works observed that training with BN could hinder performance due to the mismatch of the BN statistics between training and testing. Group Normalization (GN) is thus more often used in FL as an alternative to BN. In this paper, we identify a more fundamental issue of BN in FL that makes BN inferior even with high-frequency communication between clients and servers. We then propose a frustratingly simple treatment, which significantly improves BN and makes it outperform GN across a wide range of FL settings. Along with this study, we also reveal an unreasonable behavior of BN in FL. We find it quite robust in the low-frequency communication regime where FL is commonly believed to degrade drastically. We hope that our study could serve as a valuable reference for future practical usage and theoretical analysis in FL.
Author context
Most prolific author: 4 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 36 comparisons
Ranked above opponent in 46% of matchups.
- ▲ beat Memoria: Hebbian Memory Architecture for H… ×4
- ▼ lost to Best Arm Identification for Stochastic Ris… ×4
- ▼ lost to InstructScene: Instruction-Driven 3D Indoo… ×4
- ▼ lost to Efficient ConvBN Blocks for Transfer Learn… ×4
- ▼ lost to MT-Ranker: Reference-free machine translat… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 36)