← ICLR 2024 leaderboard

Detecting Language Model Attacks With Perplexity

Gabriel Alon, Michael J Kamfonas

generative modelsperplexityjailbreakllmadversarialattackChatGPTBARDLLaMA-2-ChatClaudegenerative modelneural networkadversarial stringadversarial suffixnlptransformer
11.80100
Fused
band ≈ ±15 pct pts (from σ = 0.30)
19.60100
Mimo
band ≈ ±20 pct pts (from σ = 0.41)
6.00100
DeepSeek
band ≈ ±23 pct pts (from σ = 0.45)

OpenReview ground truth

Rejected

TL;DR — We show that adversarial suffix attacks on LLMs can be detected with perplexity and sequence length. False positives are investigated and handled in the model.

Abstract

A novel hack involving Large Language Models (LLMs) has emerged, exploiting adversarial suffixes to deceive models into generating perilous responses. Such jailbreaks can trick LLMs into providing intricate instructions to a malicious user for creating explosives, orchestrating a bank heist, or facilitating the creation of offensive content. By evaluating the perplexity of queries with adversarial suffixes using an open-source LLM (GPT-2), we found that they have exceedingly high perplexity values. As we explored a broad range of regular (non-adversarial) prompt varieties, we concluded that false positives are a significant challenge for plain perplexity filtering. A Light-GBM trained on perplexity and token length resolved the false positives and correctly detected most adversarial attacks in the test set.

Author context

Most prolific author: 1 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Battle history — 34 comparisons

Ranked above opponent in 42% of matchups.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 34)