PapersWithELO
← ICLR 2024 leaderboard

Towards Faithful Neural Network Intrinsic Interpretation with Shapley Additive Self-Attribution

Ying Sun, Hengshu Zhu, Hui Xiong

general MLinterpretabilityneural networksadditive attributionShapley value
62.30100
Fused
band ≈ ±16 pct pts (from σ = 0.33)
61.70100
Mimo
band ≈ ±24 pct pts (from σ = 0.47)
64.20100
DeepSeek
band ≈ ±23 pct pts (from σ = 0.45)

OpenReview ground truth

Rejected

Abstract

Self-interpreting neural networks have garnered significant interest in research. Existing works in this domain often (1) lack a solid theoretical foundation ensuring genuine interpretability or (2) compromise model expressiveness. In response, we formulate a generic Additive Self-Attribution (ASA) framework. Observing the absence of Shapley value in Additive Self-Attribution, we propose Shapley Additive Self-Attributing Neural Network (SASANet), with theoretical guarantees for the self-attribution value equal to the output's Shapley values. Specifically, SASANet uses a marginal contribution-based sequential schema and internal distillation-based training strategies to model meaningful outputs for any number of features, resulting in un-approximated meaningful value function. Our experimental results indicate SASANet surpasses existing self-attributing models in performance and rivals black-box models. Moreover, SASANet is shown more precise and efficient than post-hoc methods in interpreting its own predictions.

Author context

Most prolific author: 6 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 28)