Faithful and Efficient Explanations for Neural Networks via Neural Tangent Kernel Surrogate Models
Andrew William Engel, Zhichao Wang, Natalie Frank, Ioana Dumitriu, Sutanay Choudhury, Anand Sarwate, Tony Chiang
OpenReview ground truth
TL;DR — We evaluate new approximate neural tangent kernel, including random projection variants that cost less to compute and store.
Abstract
A recent trend in explainable AI research has focused on surrogate modeling, where neural networks are approximated as simpler ML algorithms such as kernel machines. A second trend has been to utilize kernel functions in various explain-by-example or data attribution tasks. In this work, we combine these two trends to analyze approximate empirical neural tangent kernels (eNTK) for data attribution. Approximation is critical for eNTK analysis due to the high computational cost to compute the eNTK. We define new approximate eNTK and perform novel analysis on how well the resulting kernel machine surrogate models correlate with the underlying neural network. We introduce two new random projection variants of approximate eNTK which allow users to tune the time and memory complexity of their calculation. We conclude that kernel machines using approximate neural tangent kernel as the kernel function are effective surrogate models, with the introduced trace NTK the most consistent performer.
Author context
Most prolific author: 1 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 34 comparisons
Ranked above opponent in 44% of matchups.
- ▲ beat The Extrapolation Power of Implicit Models ×6
- ▼ lost to Improving length generalization in transfo… ×4
- ▼ lost to Understanding Reconstruction Attacks with … ×4
- ▼ lost to Accurate and Scalable Estimation of Episte… ×4
- ▼ lost to GOAt: Explaining Graph Neural Networks via… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 34)