← ICLR 2024 leaderboard

OpenReviewer: Mitigating Challenges in LLM Reviewing

Keith Tyser, Jason Lee, Avi Shporer, Madeleine Udell, teeni@tauex.tau.ac.il, Iddo Drori

datasets & benchmarksChatGPTreviewing research papersLLM watermarkinglong context windowsfine-tuningrandomized human blind evaluation
5.60100
Fused
band ≈ ±15 pct pts (from σ = 0.30)
8.00100
Mimo
band ≈ ±21 pct pts (from σ = 0.41)
5.20100
DeepSeek
band ≈ ±22 pct pts (from σ = 0.44)

OpenReview ground truth

Rejected

TL;DR — Using LLMs for reviewing research papers has benefits but also challenges that we address by LLM watermarking, relevant context, error and shortcoming detection, blind human evaluation, and improving quality of papers and reviews with human feedback.

Abstract

Human reviews of research papers are slow and of variable quality. Hence there is increasing interest in using large language models (LLMs) such as GPT to review research papers. This paper develops a proof-of-concept LLM review process that shows LLMs offer consistently high-quality reviews almost instantly. However, many challenges and limitations remain: risk of misuse, inflated review scores, overconfident ratings, skewed score distributions, and limited prompt length. We mitigate these issues without prompt engineering by using LLM watermarking to mark LLM-generated reviews; classifying and detection errors and shortcomings of papers; and using long-context windows that include the review form, entire paper, reviewer guidelines, code of ethics and conduct, area chair guidelies, and previous year statistics; and a blind human evaluation of reviews. We aim to use OpenReviewer to review and revise research papers, improving their quality. This work identifies and addresses drawbacks associated with GPT as a reviewer and enhances the quality of the reviewing process based on a randomized human blind evaluation. Making OpenReviewer available as an open online service that generates reviews will allow the use of scalable human feedback to learn and improve.

Author context

Most prolific author: 3 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Battle history — 32 comparisons

Ranked above opponent in 32% of matchups.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 32)