PapersWithELO
← ICLR 2024 leaderboard

A Backdoor-based Explainable AI Benchmark for Improved Fidelity in Evaluating Attribution Methods

Peiyu Yang, NAVEED AKHTAR, Jiantong Jiang, Ajmal Saeed Mian

interpretability & vizFeature AttributionExplainable AI BenchmarkBackdoor
78.70100
Fused
band ≈ ±16 pct pts (from σ = 0.31)
79.20100
Mimo
band ≈ ±23 pct pts (from σ = 0.46)
78.60100
DeepSeek
band ≈ ±21 pct pts (from σ = 0.42)

OpenReview ground truth

Rejected

TL;DR — A backdoor-based explainable AI benchmark is proposed for improved fidelity in evaluating attribution methods.

Abstract

Attribution methods compute importance scores for input features to explain the output predictions of deep models. However, accurate assessment of the performance of attribution methods is challenged by the lack of ground truth along with other confounding factors such as attribution post-processing and explanation objectives. In this paper, we first identify a set of fidelity criteria that must be satisfied for reliable evaluation of attribution methods. Then, we introduce a Trojaned model based benchmarking framework that adheres to the desired fidelity criteria. We theoretically establish the superiority of our approach over existing benchmarks for well-founded attribution evaluation. With extensive analysis, we also identify a setup for a consistent and fair benchmarking of attribution methods across different underlying methodologies. This setup is ultimately employed for a comprehensive comparison of existing methods using our benchmark. Finally, our analysis also provides guidance for defending against backdoor attacks using existing attribution methods.

Author context

Most prolific author: 3 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Battle history — 32 comparisons

Ranked above opponent in 59% of matchups.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 32)