OpenReviewer: Mitigating Challenges in LLM Reviewing
Keith Tyser, Jason Lee, Avi Shporer, Madeleine Udell, teeni@tauex.tau.ac.il, Iddo Drori
OpenReview ground truth
TL;DR — Using LLMs for reviewing research papers has benefits but also challenges that we address by LLM watermarking, relevant context, error and shortcoming detection, blind human evaluation, and improving quality of papers and reviews with human feedback.
Abstract
Human reviews of research papers are slow and of variable quality. Hence there is increasing interest in using large language models (LLMs) such as GPT to review research papers. This paper develops a proof-of-concept LLM review process that shows LLMs offer consistently high-quality reviews almost instantly. However, many challenges and limitations remain: risk of misuse, inflated review scores, overconfident ratings, skewed score distributions, and limited prompt length. We mitigate these issues without prompt engineering by using LLM watermarking to mark LLM-generated reviews; classifying and detection errors and shortcomings of papers; and using long-context windows that include the review form, entire paper, reviewer guidelines, code of ethics and conduct, area chair guidelies, and previous year statistics; and a blind human evaluation of reviews. We aim to use OpenReviewer to review and revise research papers, improving their quality. This work identifies and addresses drawbacks associated with GPT as a reviewer and enhances the quality of the reviewing process based on a randomized human blind evaluation. Making OpenReviewer available as an open online service that generates reviews will allow the use of scalable human feedback to learn and improve.
Author context
Most prolific author: 3 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 32 comparisons
Ranked above opponent in 32% of matchups.
- ▼ lost to RedMotion: Motion Prediction via Redundanc… ×6
- ▼ lost to SOTOPIA: Interactive Evaluation for Social… ×4
- ▼ lost to SKILL-MIX: a Flexible and Expandable Famil… ×4
- ▼ lost to What Makes ImageNet Look Unlike LAION ×4
- ▼ lost to CompA: Addressing the Gap in Compositional… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 32)