Detecting Language Model Attacks With Perplexity
Gabriel Alon, Michael J Kamfonas
OpenReview ground truth
TL;DR — We show that adversarial suffix attacks on LLMs can be detected with perplexity and sequence length. False positives are investigated and handled in the model.
Abstract
A novel hack involving Large Language Models (LLMs) has emerged, exploiting adversarial suffixes to deceive models into generating perilous responses. Such jailbreaks can trick LLMs into providing intricate instructions to a malicious user for creating explosives, orchestrating a bank heist, or facilitating the creation of offensive content. By evaluating the perplexity of queries with adversarial suffixes using an open-source LLM (GPT-2), we found that they have exceedingly high perplexity values. As we explored a broad range of regular (non-adversarial) prompt varieties, we concluded that false positives are a significant challenge for plain perplexity filtering. A Light-GBM trained on perplexity and token length resolved the false positives and correctly detected most adversarial attacks in the test set.
Author context
Most prolific author: 1 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 34 comparisons
Ranked above opponent in 42% of matchups.
- ▼ lost to Generative Modeling with Phase Stochastic … ×6
- ▼ lost to SAN: Inducing Metrizability of GAN with Di… ×4
- ▼ lost to Stay on Topic with Classifier-Free Guidanc… ×4
- ▼ lost to Emu: Generative Pretraining in Multimodali… ×4
- ▼ lost to Unsupervised open-vocabulary action recogn… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 34)