PapersWithELO
← ICLR 2024 leaderboard

Understanding Reconstruction Attacks with the Neural Tangent Kernel and Dataset Distillation

Noel Loo, Ramin Hasani, Mathias Lechner, Alexander Amini, Daniela Rus

general MLDataset DistillationReconstruction AttacksNeural Tangent Kernel
82.40100
Fused
band ≈ ±14 pct pts (from σ = 0.28)
80.30100
Mimo
band ≈ ±20 pct pts (from σ = 0.41)
82.80100
DeepSeek
band ≈ ±20 pct pts (from σ = 0.40)

OpenReview ground truth

Accepted

TL;DR — We analyze parameter-based reconstruction attacks from an NTK perspective and show that it is a variant of dataset distillation

Abstract

Modern deep learning requires large volumes of data, which could contain sensitive or private information that cannot be leaked. Recent work has shown for homogeneous neural networks a large portion of this training data could be reconstructed with only access to the trained network parameters. While the attack was shown to work empirically, there exists little formal understanding of its effective regime and which datapoints are susceptible to reconstruction. In this work, we first build a stronger version of the dataset reconstruction attack and show how it can provably recover the \emph{entire training set} in the infinite width regime. We then empirically study the characteristics of this attack on two-layer networks and reveal that its success heavily depends on deviations from the frozen infinite-width Neural Tangent Kernel limit. Next, we study the nature of easily-reconstructed images. We show that both theoretically and empirically, reconstructed images tend to ``outliers'' in the dataset, and that these reconstruction attacks can be used for \textit{dataset distillation}, that is, we can retrain on reconstructed images and obtain high predictive accuracy.

Author context

Most prolific author: 5 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Battle history — 38 comparisons

Ranked above opponent in 59% of matchups.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 38)