PapersWithELO
← ICLR 2024 leaderboard

Improving Robustness and Accuracy with Retrospective Online Adversarial Distillation

Joongsu Kim, Junhyung Jo, Suha Kwak, Young-Joo Suh

fairness, safety & privacyAdversarial TrainingAdversarial DistillationKnowledge Distillation
63.50100
Fused
band ≈ ±14 pct pts (from σ = 0.29)
53.10100
Mimo
band ≈ ±21 pct pts (from σ = 0.42)
72.80100
DeepSeek
band ≈ ±20 pct pts (from σ = 0.40)

OpenReview ground truth

Rejected

TL;DR — We propose retrospective online adversarial distillation (ROAD), to improve robustness against adversarial attacks and natural accuracy.

Abstract

Adversarial distillation (AD), transferring knowledge of a robust teacher model to a student model, has emerged as an advanced technique for improving robustness against adversarial attacks. However, AD in general suffers from the high computational complexity of pre-training the robust teacher, and the inherent trade-off between robustness and natural accuracy (i.e., accuracy on clean data). To address these issues, we propose retrospective online adversarial distillation (ROAD). ROAD exploits the student itself of the last epoch and a natural model (i.e., a model trained with clean data) as teachers, instead of a pre-trained robust teacher in the conventional AD. We revealed both theoretically and empirically that knowledge distillation from the student of the last epoch allows to penalize overly confident predictions on adversarial examples, leading to improved robustness and generalization. Also, the student and the natural model are trained together in a collaborative manner, which enables to improve natural accuracy of the student more effectively. We demonstrate by extensive experiments that ROAD achieved outstanding performance in both robustness and natural accuracy with substantially reduced training time and computation cost.

Author context

Most prolific author: 3 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 34)