Human-in-the-loop Detection of AI-generated Text via Grammatical Patterns
Johan Lokna, Mislav Balunovic, Martin Vechev
OpenReview ground truth
Abstract
The increasing proliferation of large language models (LLMs) has raised significant concerns about the detection of AI-written text. Ideally, the detection method should be accurate (in particular, it should not falsely accuse humans of using AI-generated text), and interpretable (it should provide a decision as to why the text was detected as either human or AI-generated). Existing methods tend to fall short of one or both of these requirements, and recent work has even shown that detection is impossible in the full generality. In this work, we focus on the problem of detecting AI-generated text in a domain where a training dataset of human-written samples is readily available. Our key insight is to learn interpretable grammatical patterns that are highly indicative of human or AI written text. The most useful of these patterns can then be given to humans as part of a human-in-the-loop approach. In our experimental evaluation, we show that the approach can effectively detect AI-written text in a variety of domains and generalize to different language models. Our results in a human trial show an improvement in the detection accuracy from $43$% to $86$%, demonstrating the effectiveness of the human-in-the-loop approach. We also show that the method is robust to different ways of prompting LLM to generate human-like patterns. Overall, our study demonstrates that AI text can be accurately and interpretably detected using a human-in-the-loop approach.
Author context
Most prolific author: 10 submissions (credibility 0.79).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 32 comparisons
Ranked above opponent in 46% of matchups.
- ▲ beat Empirical Likelihood for Fair Classificati… ×6
- ▲ beat Rethinking the Buyer’s Inspection Paradox … ×4
- ▼ lost to User Inference Attacks on Large Language M… ×4
- ▼ lost to Negative Label Guided OOD Detection with P… ×4
- ▼ lost to Fair Classifiers that Abstain without Harm ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 32)