Safeguarding Data in Multimodal AI: A Differentially Private Approach to CLIP Training
Alyssa Huang, Peihan Liu, Ryumei Nakada, Linjun Zhang, Wanrong Zhang
OpenReview ground truth
Abstract
The surge in multimodal AI's success has sparked concerns over data privacy in vision-and-language tasks. While CLIP has revolutionized multimodal learning through joint training on images and text, its potential to unintentionally disclose sensitive information necessitates the integration of privacy-preserving mechanisms. We introduce a differentially private adaptation of the Contrastive Language-Image Pretraining (CLIP) model that effectively addresses privacy concerns while retaining accuracy. Our proposed method, \dpclip, is rigorously evaluated on benchmark datasets encompassing diverse vision-and-language tasks such as image classification and image captioning. We demonstrate that our approach retains performance on par with the standard non-private CLIP model. Furthermore, we analyze our proposed algorithm under linear representation settings. We derive the convergence rate of our algorithm and show a trade-off between utility and privacy when gradients are clipped per-\textit{batch} and the loss function does not satisfy smoothness conditions assumed in the literature for the analysis of DP-SGD.
Author context
Most prolific author: 5 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 36 comparisons
Ranked above opponent in 35% of matchups.
- ▲ beat NDIM: Neuronal Diversity Inspired Model fo… ×8
- ▲ beat Dispatching Ambulances using Deep Reinforc… ×6
- ▲ beat OWL: A Large Language Model for IT Operati… ×6
- ▼ lost to Rethinking Self-Supervise Learning: An Ins… ×4
- ▼ lost to On robust overfitting: adversarial trainin… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 36)