Calibration Attack: A Framework For Adversarial Attacks Targeting Calibration
Stephen Obadinma, Xiaodan Zhu, Hongyu Guo
OpenReview ground truth
TL;DR — We develop a new framework for using adversarial attacks to prioritize damaging calibration over accuracy and show the potential severity that these attacks have.
Abstract
We introduce a new framework of adversarial attacks, named calibration attacks, in which the attacks are generated and organized to trap victim models to be miscalibrated without altering their original accuracy, hence seriously endangering the trustworthiness of the models and any decision-making based on their confidence scores. Specifically, we identify four novel forms of calibration attacks: underconfidence attacks, overconfidence attacks, maximum miscalibration attacks, and random confidence attacks, in both the black-box and white-box setups. We then test these new attacks on typical victim models with comprehensive datasets, demonstrating that even with a relatively low number of queries, the attacks can create significant calibration mistakes. We further provide detailed analyses to understand different aspects of calibration attacks. Building on that, we investigate the effectiveness of widely used adversarial defences and calibration methods against these types of attacks, which then inspires us to devise two novel defences against such calibration attacks.
Author context
Most prolific author: 4 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 42 comparisons
Ranked above opponent in 47% of matchups.
- ▼ lost to Test like you Train in Implicit Deep Learn… ×6
- ▲ beat Does Writing with Language Models Reduce C… ×6
- ▼ lost to Multiobjective Stochastic Linear Bandits u… ×4
- ▲ beat OmniControl: Control Any Joint at Any Time… ×4
- ▲ beat WASA: WAtermark-based Source Attribution f… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 42)