Sparse Backpropagation for MoE Training
Liyuan Liu, Jianfeng Gao, Weizhu Chen
OpenReview ground truth
TL;DR — We introduce SparseMixer, a scalable gradient estimator that bridges the gap between backpropagation and sparse expert routing.
Abstract
One defining characteristic of Mixture-of-Expert (MoE) models is their capacity for conducting sparse computation via expert routing, leading to remarkable scalability. However, backpropagation, the cornerstone of deep learning, requires dense computation, thereby posting challenges in MoE gradient computations. Here, we introduce SparseMixer, a scalable gradient estimator that bridges the gap between backpropagation and sparse expert routing. Unlike typical MoE training which strategically neglects certain gradient terms for the sake of sparse computation and scalability, SparseMixer provides scalable gradient approximations for these terms, enabling reliable gradient estimation in MoE training. Grounded in a numerical ODE framework, SparseMixer harnesses the mid-point method, a second-order ODE solver, to deliver precise gradient approximations with negligible computational overhead. Applying SparseMixer to Switch Transformer on both pre-training and machine translation tasks, SparseMixer showcases considerable performance gain, accelerating training convergence by up to 2 times.
Author context
Most prolific author: 13 submissions (credibility 0.86).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 40 comparisons
Ranked above opponent in 56% of matchups.
- ▼ lost to Zero-Shot Robotic Manipulation with Pre-Tr… ×8
- ▲ beat SciRE-Solver: Accelerating Diffusion Model… ×6
- ▼ lost to Language Model Decoding as Direct Metrics … ×6
- ▼ lost to Improved Active Learning via Dependent Lev… ×6
- ▲ beat Constructing Sparse Neural Architecture wi… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 40)