LogicBench: Towards Systematic Evaluation of Logical Reasoning Ability of Large Language Models
Mihir Parmar, Neeraj Varshney, Nisarg Patel, Man Luo, Santosh Mashetty, Arindam Mitra, Chitta Baral
OpenReview ground truth
TL;DR — We present LogicBench, a systematically created question-answering natural language dataset for logical reasoning, encompassing 25 reasoning patterns spanning across three logic types, namely, propositional, first-order, and non-monotonic logics.
Abstract
Recently developed large language models (LLMs) have been shown to perform remarkably well on a wide range of language understanding tasks. But, can they really "reason" over the natural language? This question has been receiving significant research attention and a number of reasoning skills such as commonsense, numerical, and qualitative have been studied. However, the crucial skill pertaining to 'logical reasoning' has remained underexplored. Existing work investigating this reasoning ability has focused only on a couple of inference rules (such as modus ponens and modus tollens) of propositional and first-order logic. To enable systematic evaluation of logical reasoning, we introduce LogicBench, a natural language question-answering dataset encompassing 25 different reasoning patterns spanning over propositional, first-order, and non-monotonic logics. Key steps of our dataset construction consist of (1) controlled generation of sentences and their negations containing different ontologies, (2) (context, question, answer) triplets creation using heuristically designed templates, and (3) semantic variations of triplets adding more diversity. We present a comprehensive evaluation with a range of LLMs such as GPT-4, GPT-3, ChatGPT, and FLAN-T5 using chain-of-thought prompting in both zero-shot and few-shot settings. Experimental results show that existing LLMs do not fare well on LogicBench; especially, they struggle on instances requiring complex reasoning steps. Furthermore, we also show that LLMs trained using our data exhibit a better understanding of logical reasoning leading to performance improvements on several existing logical reasoning datasets such as LogicNLI, FOLIO, LogiQA, and ReClor.
Author context
Most prolific author: 2 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 34 comparisons
Ranked above opponent in 57% of matchups.
- ▼ lost to SWE-bench: Can Language Models Resolve Rea… ×4
- ▲ beat Stochastic Subnetwork Annealing: A Regular… ×4
- ▼ lost to FLASK: Fine-grained Language Model Evaluat… ×4
- ▲ beat What, when, and where? -- Self-Supervised … ×4
- ▼ lost to Instance Needs More Care: Rewriting Prompt… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 34)