PapersWithELO
← ICLR 2024 leaderboard

Improving Private Training via In-distribution Public Data Synthesis and Generalization

Jinseong Park, Yujin Choi, Jaewook Lee

fairness, safety & privacyDifferential PrivacyPrivacyOptimizationDP-SGDDiffusionSynthesis
29.90100
Fused
band ≈ ±13 pct pts (from σ = 0.27)
27.90100
Mimo
band ≈ ±19 pct pts (from σ = 0.39)
29.40100
DeepSeek
band ≈ ±19 pct pts (from σ = 0.37)

OpenReview ground truth

Rejected

TL;DR — Enhancing differentially private classification through in-distribution public data using diffusion synthesis and optimization for well-generalizing minima.

Abstract

To alleviate the utility degradation of deep learning classification with differential privacy (DP), employing extra public data or pre-trained models has been widely explored. Recently, the use of in-distribution public data has been investigated, where a tiny subset of data owners share their data publicly. In this paper, to mitigate memorization and overfitting by the limited-sized in-distribution public data, we leverage recent diffusion models and employ various augmentation techniques for improving diversity. We then explore the optimization to discover flat minima to public data and suggest weight multiplicity to enhance the generalization of the private training. While assuming 4\% of training data as public, our method brings significant performance gain even without using pre-trained models, i.e., achieving 85.78\% on CIFAR-10 with a privacy budget of $\varepsilon=2$ and $\delta=10^{-5}$.

Author context

Most prolific author: 1 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 42)