PapersWithELO
← ICLR 2024 leaderboard

CARSO: Blending Adversarial Training and Purification Improves Adversarial Robustness

Emanuele Ballarin, Alessio ansuini, Luca Bortolussi

fairness, safety & privacyadversarial robustnessadversarial trainingadversarial purificationgenerative purificationinternal representation
87.80100
Fused
band ≈ ±14 pct pts (from σ = 0.27)
86.20100
Mimo
band ≈ ±20 pct pts (from σ = 0.40)
91.20100
DeepSeek
band ≈ ±19 pct pts (from σ = 0.37)

OpenReview ground truth

Rejected

TL;DR — We propose a novel adversarial defence for image classifiers, merging adversarial training and purification: the internal representation of an adversarially-trained classifier is mapped to a distribution of denoised reconstructions to be classified.

Abstract

In this work, we propose a novel adversarial defence mechanism for image classification - CARSO - blending the paradigms of *adversarial training* and *adversarial purification* in a mutually-beneficial, robustness-enhancing way. The method builds upon an adversarially-trained classifier, and learns to map its *internal representation* associated with a potentially perturbed input onto a distribution of tentative clean reconstructions. Multiple samples from such distribution are classified by the adversarially-trained model, and an aggregation of its outputs finally constitutes the *robust prediction* of interest. Experimental evaluation by a well-established benchmark of varied, strong adaptive attacks, across different image datasets and classifier architectures, shows that CARSO is able to defend itself against foreseen and unforeseen threats, including adaptive *end-to-end* attacks devised for stochastic defences. Paying a tolerable *clean* accuracy toll, our method improves by a significant margin the *state of the art* for CIFAR-10 and CIFAR-100 $\ell_\infty$ robust classification accuracy against AutoAttack.

Author context

Most prolific author: 2 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 44)