Generative Models are Self-Watermarked: Intellectual Property Declaration through Re-Generation
Aditya Desu, Xuanli He, Qiongkai Xu, Wei Lu
OpenReview ground truth
TL;DR — Our work discovers the intrinsic fingerprint in the generative model and proposes an iterative re-generation approach to reinforce such a signal.
Abstract
Protecting intellectual property for generated data has emerged as a critical concern for AI corporations, as machine-generated content proliferates. Reusing generated data without permission poses a formidable barrier to safeguarding the intellectual property tied to these models. The verification of data ownership is further complicated by the use of Machine Learning as a Service (MLaaS), which often operates as a black-box system. Our work is dedicated to detecting data reuse from even an individual sample. In contrast to watermarking techniques that embed additional information as watermark triggers into models or generated content, our approach does not introduce artificial watermarks which may compromise the quality of model outputs. Our investigation reveals the existence of latent fingerprints inherently present within deep learning models. In response, we propose an explainable verification procedure to verify data ownership through re-generation. Furthermore, we introduce a novel methodology to amplify the model fingerprints through iterative data regeneration and a theoretical grounding on the proposed approach. We demonstrate the viability of our approach using recent advanced text and image generative models.
Author context
Most prolific author: 2 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 40 comparisons
Ranked above opponent in 51% of matchups.
- ▲ beat Advancing Test-Time Adaptation for Acousti… ×6
- ▲ beat FLAT-Chat: A Word Recovery Attack on Feder… ×6
- ▼ lost to Scalabale AI Safety via Doubly-Efficient D… ×4
- ▼ lost to Scaling up Trustless DNN Inference with Ze… ×4
- ▲ beat Diffusion Denoising as a Certified Defense… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 40)