PapersWithELO
← ICLR 2024 leaderboard

Safe and Robust Watermark Injection with a Single OoD Image

Shuyang Yu, Junyuan Hong, Haobo Zhang, Haotao Wang, Zhangyang Wang, Jiayu Zhou

fairness, safety & privacyBackdoorWatermarkingrobustness
35.80100
Fused
band ≈ ±14 pct pts (from σ = 0.28)
34.20100
Mimo
band ≈ ±20 pct pts (from σ = 0.40)
31.60100
DeepSeek
band ≈ ±20 pct pts (from σ = 0.40)

OpenReview ground truth

Accepted

Abstract

Training a high-performance deep neural network requires large amounts of data and computational resources. Protecting the intellectual property (IP) and commercial ownership of a deep model is challenging yet increasingly crucial. A major stream of watermarking strategies implants verifiable backdoor triggers by poisoning training samples, but these are often unrealistic due to data privacy and safety concerns and are vulnerable to minor model changes such as fine-tuning. To overcome these challenges, we propose a safe and robust backdoor-based watermark injection technique that leverages the diverse knowledge from a single out-of-distribution (OoD) image, which serves as a secret key for IP verification. The independence of training data makes it agnostic to third-party promises of IP security. We induce robustness via random perturbation of model parameters during watermark injection to defend against common watermark removal attacks, including fine-tuning, pruning, and model extraction. Our experimental results demonstrate that the proposed watermarking approach is not only time- and sample-efficient without training data, but also robust against the watermark removal attacks above.

Author context

Most prolific author: 24 submissions (credibility 0.23).

Delta if applied: -1.1 percentile

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 36)