PapersWithELO
← ICLR 2024 leaderboard

A Change of Heart: Backdoor Attacks on Security-Centric Diffusion Models

Changjiang Li, Ren Pang, Bochuan Cao, Jinghui Chen, Ting Wang

fairness, safety & privacyDiffusion ModelsAdversarial PurificationBackdoor Attacks
91.20100
Fused
band ≈ ±15 pct pts (from σ = 0.29)
94.40100
Mimo
band ≈ ±21 pct pts (from σ = 0.42)
88.90100
DeepSeek
band ≈ ±20 pct pts (from σ = 0.41)

OpenReview ground truth

Rejected

TL;DR — DIFF2 is a backdoor attack on diffusion models, highlighting risks in their use for critical applications (e.g., adversarial purification and robustness certification)

Abstract

Diffusion models have been employed as defensive tools to reinforce the security of other models, notably in purifying adversarial examples and certifying adversarial robustness. Meanwhile, the prohibitive training costs often make the use of pre-trained diffusion models an attractive practice. The tension between the intended use of these models and their unvalidated nature raises significant security concerns that remain largely unexplored. To bridge this gap, we present DIFF2, a novel backdoor attack tailored to security-centric diffusion models. Essentially, DIFF2 superimposes a diffusion model with a malicious diffusion-denoising process, guiding inputs embedded with specific triggers toward an adversary-defined distribution, while preserving the normal process for other inputs. Our case studies on adversarial purification and robustness certification show that DIFF2 substantially diminishes both post-purification and certified accuracy across various benchmark datasets and diffusion models, highlighting the potential risks of utilizing pre-trained diffusion models as defensive tools. We further explore possible countermeasures, suggesting promising avenues for future research.

Author context

Most prolific author: 7 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 38)