Amplifying Training Data Exposure through Fine-Tuning with Pseudo-Labeled Memberships
Myung Gyo Oh, Hong Eun Ahn, Leo Hyun Park, Taekyoung Kwon
OpenReview ground truth
TL;DR — We present a novel attack method that amplifies training data exposure in language models by fine-tuning them with pseudo-labeled memberships.
Abstract
Neural language models (LMs) are vulnerable to training data extraction attacks due to data memorization. This paper introduces a novel attack scenario wherein an attacker adversarially fine-tunes pre-trained LMs to amplify the exposure of the original training data. This strategy differs from prior studies by aiming to intensify the LM's retention of its pre-training dataset. To achieve this, the attacker needs to collect generated texts that are closely aligned with the pre-training data. However, without knowledge of the actual dataset, quantifying the amount of pre-training data within generated texts is challenging. To address this, we propose the use of pseudo-labels for these generated texts, leveraging membership approximations indicated by machine-generated probabilities from the target LM. We subsequently fine-tune the LM to favor generations with higher likelihoods of originating from the pre-training data, based on their membership probabilities. Our empirical findings indicate a remarkable outcome: LMs with over 1B parameters exhibit a four to eight-fold increase in training data exposure. We discuss potential mitigations and suggest future research directions.
Author context
Most prolific author: 1 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 46 comparisons
Ranked above opponent in 49% of matchups.
- ▲ beat Model Inversion Robustness: Can Transfer L… ×6
- ▲ beat Structured Pruning Adapters ×6
- ▲ beat Towards the Vulnerability of Watermarking … ×4
- ▲ beat A Comprehensive Study of Privacy Risks in … ×4
- ▼ lost to STanHop: Sparse Tandem Hopfield Model for … ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 46)