Improving Private Training via In-distribution Public Data Synthesis and Generalization
Jinseong Park, Yujin Choi, Jaewook Lee
OpenReview ground truth
TL;DR — Enhancing differentially private classification through in-distribution public data using diffusion synthesis and optimization for well-generalizing minima.
Abstract
To alleviate the utility degradation of deep learning classification with differential privacy (DP), employing extra public data or pre-trained models has been widely explored. Recently, the use of in-distribution public data has been investigated, where a tiny subset of data owners share their data publicly. In this paper, to mitigate memorization and overfitting by the limited-sized in-distribution public data, we leverage recent diffusion models and employ various augmentation techniques for improving diversity. We then explore the optimization to discover flat minima to public data and suggest weight multiplicity to enhance the generalization of the private training. While assuming 4\% of training data as public, our method brings significant performance gain even without using pre-trained models, i.e., achieving 85.78\% on CIFAR-10 with a privacy budget of $\varepsilon=2$ and $\delta=10^{-5}$.
Author context
Most prolific author: 1 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 42 comparisons
Ranked above opponent in 44% of matchups.
- ▲ beat Detecting Shortcuts using Mutual Informati… ×6
- ▼ lost to Scalabale AI Safety via Doubly-Efficient D… ×4
- ▼ lost to Scaling up Trustless DNN Inference with Ze… ×4
- ▲ beat Culture in Artificial Intelligence: A Lite… ×4
- ▼ lost to ConjNorm: Tractable Density Estimation for… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 42)