A Backdoor-based Explainable AI Benchmark for Improved Fidelity in Evaluating Attribution Methods
Peiyu Yang, NAVEED AKHTAR, Jiantong Jiang, Ajmal Saeed Mian
OpenReview ground truth
TL;DR — A backdoor-based explainable AI benchmark is proposed for improved fidelity in evaluating attribution methods.
Abstract
Attribution methods compute importance scores for input features to explain the output predictions of deep models. However, accurate assessment of the performance of attribution methods is challenged by the lack of ground truth along with other confounding factors such as attribution post-processing and explanation objectives. In this paper, we first identify a set of fidelity criteria that must be satisfied for reliable evaluation of attribution methods. Then, we introduce a Trojaned model based benchmarking framework that adheres to the desired fidelity criteria. We theoretically establish the superiority of our approach over existing benchmarks for well-founded attribution evaluation. With extensive analysis, we also identify a setup for a consistent and fair benchmarking of attribution methods across different underlying methodologies. This setup is ultimately employed for a comprehensive comparison of existing methods using our benchmark. Finally, our analysis also provides guidance for defending against backdoor attacks using existing attribution methods.
Author context
Most prolific author: 3 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 32 comparisons
Ranked above opponent in 59% of matchups.
- ▲ beat Understanding Deep Neural Networks as Dyna… ×4
- ▲ beat Information based explanation methods for … ×4
- ▲ beat Conditional MAE: An Empirical Study of Mul… ×4
- ▲ beat On the Joint Interaction of Models, Data, … ×4
- ▲ beat Structured Pruning Adapters ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 32)