Can LLM-Generated Misinformation Be Detected?
Canyu Chen, Kai Shu
OpenReview ground truth
TL;DR — We discover that LLM-generated misinformation can be harder to detect for humans and detectors compared to human-written misinformation with the same semantics, which suggests it can have more deceptive styles and potentially cause more harm.
Abstract
The advent of Large Language Models (LLMs) has made a transformative impact. However, the potential that LLMs such as ChatGPT can be exploited to generate misinformation has posed a serious concern to online safety and public trust. A fundamental research question is: will LLM-generated misinformation cause more harm than human-written misinformation? We propose to tackle this question from the perspective of detection difficulty. We first build a taxonomy of LLM-generated misinformation. Then we categorize and validate the potential real-world methods for generating misinformation with LLMs. Then, through extensive empirical investigation, we discover that LLM-generated misinformation can be harder to detect for humans and detectors compared to human-written misinformation with the same semantics, which suggests it can have more deceptive styles and potentially cause more harm. We also discuss the implications of our discovery on combating misinformation in the age of LLMs and the countermeasures.
Author context
Most prolific author: 2 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 32 comparisons
Ranked above opponent in 43% of matchups.
- ▼ lost to Large Scene Synthesis Controlled With Deta… ×8
- ▲ beat Sparse Model Soups: A Recipe for Improved … ×6
- ▼ lost to Certifying LLM Safety against Adversarial … ×4
- ▼ lost to Detecting Pretraining Data from Large Lang… ×4
- ▲ beat Rethinking the Buyer’s Inspection Paradox … ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 32)