PapersWithELO
← ICLR 2024 leaderboard

From Malicious to Marvelous: The Art of Adversarial Attack as Diffusion

Guangrun Wang, Chen Lin, Philip Torr

general MLAdversarial RobustnessCIFAR-10Reverse Adversarial ProcessDiffusion Model
58.30100
Fused
band ≈ ±17 pct pts (from σ = 0.33)
30.00100
Mimo
band ≈ ±24 pct pts (from σ = 0.47)
87.90100
DeepSeek
band ≈ ±23 pct pts (from σ = 0.47)

OpenReview ground truth

Rejected

Abstract

The ubiquitous presence of adversarial attacks in deep learning has been a source of frustration and challenge for researchers for years. However, in this work, we establish a new connection between adversarial attacks and the intricate process of diffusion. Specifically, we formulate an adversarial attack as a diffusion process, and by reverting this adversarial attack process, we have devised an innovative defense mechanism that stands out as a general-purpose defense against both black-box and white-box attacks. We call this new mechanism a Reverse Adversarial Process (RAP), which is ensured by a theoretical treatment for deploying denoising diffusion models on arbitrary distributions. Empirically, we found our model successfully defends against adversarial attacks with an unprecedented level of accuracy. For example, our approach has demonstrated exceptional performance on the \textit{RobustBench}, a highly-regarded leaderboard for assessing adversarial robustness, outperforming previous state-of-the-art methods by a clear margin.

Author context

Most prolific author: 23 submissions (credibility 0.20).

Delta if applied: -1.2 percentile

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 32)