PapersWithELO
← ICLR 2024 leaderboard

Calibration Attack: A Framework For Adversarial Attacks Targeting Calibration

Stephen Obadinma, Xiaodan Zhu, Hongyu Guo

fairness, safety & privacyrobustnesscalibrationdeep learningimage classificationadversarial
70.10100
Fused
band ≈ ±13 pct pts (from σ = 0.27)
70.80100
Mimo
band ≈ ±19 pct pts (from σ = 0.39)
69.30100
DeepSeek
band ≈ ±19 pct pts (from σ = 0.37)

OpenReview ground truth

Rejected

TL;DR — We develop a new framework for using adversarial attacks to prioritize damaging calibration over accuracy and show the potential severity that these attacks have.

Abstract

We introduce a new framework of adversarial attacks, named calibration attacks, in which the attacks are generated and organized to trap victim models to be miscalibrated without altering their original accuracy, hence seriously endangering the trustworthiness of the models and any decision-making based on their confidence scores. Specifically, we identify four novel forms of calibration attacks: underconfidence attacks, overconfidence attacks, maximum miscalibration attacks, and random confidence attacks, in both the black-box and white-box setups. We then test these new attacks on typical victim models with comprehensive datasets, demonstrating that even with a relatively low number of queries, the attacks can create significant calibration mistakes. We further provide detailed analyses to understand different aspects of calibration attacks. Building on that, we investigate the effectiveness of widely used adversarial defences and calibration methods against these types of attacks, which then inspires us to devise two novel defences against such calibration attacks.

Author context

Most prolific author: 4 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 42)