PapersWithELO
← ICLR 2024 leaderboard

LegoMT2: Non-Blocking Federated Learning for Massive Multilingual Machine Translation

Fei Yuan, Yinquan Lu, Lingpeng Kong, Lei Li, Jingjing Xu

general MLMassively Multilingual Machine TranslationNon-blockingFederated Learning
68.70100
Fused
band ≈ ±15 pct pts (from σ = 0.31)
62.40100
Mimo
band ≈ ±22 pct pts (from σ = 0.43)
75.70100
DeepSeek
band ≈ ±22 pct pts (from σ = 0.44)

OpenReview ground truth

Rejected

Abstract

What is the maximal number of languages that a single machine translation model can translate? It is a critical challenge to learn a single model for massive languages. Prior methods focus on increasing the model size and training data size. However, large models are difficult to optimize efficiently even with distributed parallel training and translation capacity can interfere among languages. To address the challenge, we propose LegoMT2, an efficient approach with a tailored model architecture for massive multilingual neural machine translation. LegoMT2 organizes 435 languages into 8 language-centric groups and attributes one local encoder-decoder for each group and a global encoder-decoder for all languages. LegoMT2 then trains each local and global encoder-decoder on a group-dedicated set of clients through asynchronous updating of parameters. We trained LegoMT2 on a large dataset with 25 billion sentence pairs beyond English-centric. LegoMT2 is 16.2$\times$ faster than the distributed training method for the same-size NLLB while improving the translation results by an average of 2.2 BLEU on \textit{Flores-101}~\footnote{We will release the model and code to the public.}.

Author context

Most prolific author: 10 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 32)