PapersWithELO
← ICLR 2024 leaderboard

FLAT-Chat: A Word Recovery Attack on Federated Language Model Training

Qiongkai Xu, Jun Wang, Olga Ohrimenko, Trevor Cohn

fairness, safety & privacyLabel inference attackLarge-scale language modelMatrix flattening
57.60100
Fused
band ≈ ±15 pct pts (from σ = 0.30)
60.90100
Mimo
band ≈ ±22 pct pts (from σ = 0.45)
48.30100
DeepSeek
band ≈ ±21 pct pts (from σ = 0.41)

OpenReview ground truth

Rejected

TL;DR — We develop a label inference attack based on gradients from federated large language model training, which can identify tokens used in training even when applied to large batches as used in modern large language models with large vocabulary sizes.

Abstract

Gradient exchange is widely applied in collaborative training of machine learning models, including Federated Learning. Curious-but-honest participants could potentially infer the output labels in recently used training data by analyzing the latest gradient updates. Previous works mostly demonstrate the attack performance under constraint training settings, such as dozens of short sentences in a batch and a small output space for labels. In this work, we propose a novel gradient flattening attack on the last linear layer of a language model, which significantly improves the attacker's efficiency in inferring the words used in training. We validate the capability of the attack on two language generation tasks: machine translation and language modeling. The attack environment is scaled up to industrial settings of a large output vocabulary and realistic training batch sizes. To mitigate the negative impact of the new attack, we explore two defense methods and demonstrate that adding differential privacy with small noise could effectively defend against our new attack without degrading model utility.

Author context

Most prolific author: 2 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 32)