PapersWithELO
← ICLR 2024 leaderboard

Amplifying Training Data Exposure through Fine-Tuning with Pseudo-Labeled Memberships

Myung Gyo Oh, Hong Eun Ahn, Leo Hyun Park, Taekyoung Kwon

fairness, safety & privacylarge language modeltraining data extractionfine-tuningpseudo-labeling with membershipprivacy
64.50100
Fused
band ≈ ±13 pct pts (from σ = 0.26)
68.70100
Mimo
band ≈ ±18 pct pts (from σ = 0.35)
60.60100
DeepSeek
band ≈ ±19 pct pts (from σ = 0.38)

OpenReview ground truth

Rejected

TL;DR — We present a novel attack method that amplifies training data exposure in language models by fine-tuning them with pseudo-labeled memberships.

Abstract

Neural language models (LMs) are vulnerable to training data extraction attacks due to data memorization. This paper introduces a novel attack scenario wherein an attacker adversarially fine-tunes pre-trained LMs to amplify the exposure of the original training data. This strategy differs from prior studies by aiming to intensify the LM's retention of its pre-training dataset. To achieve this, the attacker needs to collect generated texts that are closely aligned with the pre-training data. However, without knowledge of the actual dataset, quantifying the amount of pre-training data within generated texts is challenging. To address this, we propose the use of pseudo-labels for these generated texts, leveraging membership approximations indicated by machine-generated probabilities from the target LM. We subsequently fine-tune the LM to favor generations with higher likelihoods of originating from the pre-training data, based on their membership probabilities. Our empirical findings indicate a remarkable outcome: LMs with over 1B parameters exhibit a four to eight-fold increase in training data exposure. We discuss potential mitigations and suggest future research directions.

Author context

Most prolific author: 1 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Battle history — 46 comparisons

Ranked above opponent in 49% of matchups.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 46)