PapersWithELO
← ICLR 2024 leaderboard

Backdoor Attack for Federated Learning with Fake Clients

Pei Fang, Bochuan Cao, Jinyuan Jia, Jinghui Chen

fairness, safety & privacyBackdoor AtttackFederated Learning
48.30100
Fused
band ≈ ±13 pct pts (from σ = 0.26)
43.80100
Mimo
band ≈ ±18 pct pts (from σ = 0.36)
55.40100
DeepSeek
band ≈ ±19 pct pts (from σ = 0.38)

OpenReview ground truth

Rejected

Abstract

Federated Learning (FL) is a popular distributed machine learning paradigm that enables joint model training without sharing clients’ data. Recent studies show that federated learning can be vulnerable to potential backdoor attacks from malicious clients: such attacks aim to mislead the global model into a targeted misprediction when a specific trigger pattern is presented. Although various types of federated backdoor attacks are proposed, most of them rely on the malicious client's local data to inject the backdoor trigger into the model. In this paper, we consider a new and more challenging scenario that the attacker can only control the fake clients, who do not possess any real data at all. Such a threat model sets a higher standard for the attacker that the attack must be conducted without relying on any real client data (only knowing the target class label). Meanwhile, the resulting malicious update should not be easily detected by the potential defenses. Specifically, we first simulate the normal client updates via modeling the historical global model trajectory. Then we simultaneously optimize the backdoor trigger and manipulate the model parameters in a data-free manner to achieve our attacking goal. Extensive experiments on multiple benchmark datasets show the effectiveness of the proposed attack in the fake client setting under state-of-the-art defenses.

Author context

Most prolific author: 10 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Battle history — 42 comparisons

Ranked above opponent in 50% of matchups.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 42)