LegoMT2: Non-Blocking Federated Learning for Massive Multilingual Machine Translation
Fei Yuan, Yinquan Lu, Lingpeng Kong, Lei Li, Jingjing Xu
OpenReview ground truth
Abstract
What is the maximal number of languages that a single machine translation model can translate? It is a critical challenge to learn a single model for massive languages. Prior methods focus on increasing the model size and training data size. However, large models are difficult to optimize efficiently even with distributed parallel training and translation capacity can interfere among languages. To address the challenge, we propose LegoMT2, an efficient approach with a tailored model architecture for massive multilingual neural machine translation. LegoMT2 organizes 435 languages into 8 language-centric groups and attributes one local encoder-decoder for each group and a global encoder-decoder for all languages. LegoMT2 then trains each local and global encoder-decoder on a group-dedicated set of clients through asynchronous updating of parameters. We trained LegoMT2 on a large dataset with 25 billion sentence pairs beyond English-centric. LegoMT2 is 16.2$\times$ faster than the distributed training method for the same-size NLLB while improving the translation results by an average of 2.2 BLEU on \textit{Flores-101}~\footnote{We will release the model and code to the public.}.
Author context
Most prolific author: 10 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 32 comparisons
Ranked above opponent in 50% of matchups.
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 32)