Large-Scale Public Data Improves Differentially Private Image Generation Quality
Ruihan Wu, Chuan Guo, Kamalika Chaudhuri
OpenReview ground truth
Abstract
Public data has been frequently used to improve the privacy-accuracy trade-off of differentially private machine learning, but prior work largely assumes that this data come from the same distribution as the private. In this work, we look at how to use *generic* large-scale public data to improve the quality of differentially private image generation in Generative Adversarial Networks (GANs), and provide an improved method that uses public data effectively. Our method works under the assumption that the support of the public data distribution contains the support of the private; an example of this is when the public data come from a general-purpose internet-scale image source, while the private data consist of images of a specific type. Detailed evaluations show that our method achieves SOTA in terms of FID score and other metrics compared with existing methods that use public data, and can generate high-quality, photo-realistic images in a differentially private manner.
Author context
Most prolific author: 5 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 40 comparisons
Ranked above opponent in 52% of matchups.
- ▲ beat A Data-Driven Measure of Relative Uncertai… ×4
- ▲ beat A Recipe for Watermarking Diffusion Models ×4
- ▼ lost to A Change of Heart: Backdoor Attacks on Sec… ×4
- ▼ lost to CARSO: Blending Adversarial Training and P… ×4
- ▲ beat Boosting Backdoor Attack with A Learnable … ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 40)