On the Paradox of Generalizable Logical Reasoning in Large Language Models
Xiaojuan Tang, Zilong Zheng, Jiaqi Li, Fanxu Meng, Song-Chun Zhu, Yitao Liang, Muhan Zhang
OpenReview ground truth
Abstract
The emergent few-shot reasoning capabilities of Large Language Models (LLMs) have excited the natural language and machine learning community over recent years. Despite the numerous successful applications, it remains an open question whether LLMs have generalizable logical reasoning abilities. In this work, we expose a surprising failure of generalization in logical reasoning tasks (deduction, induction, and abduction)---when semantics are decoupled from the language reasoning process (\ie, replacing semantic words with pure symbols), LLMs tend to perform much worse. We hypothesize that the learned \textit{semantics} of language tokens do the most heavy lifting during the reasoning process but fail to imitate the basic formal reasoning abilities of humans. Furthermore, we also attempt to fine-tune Llama-2 on pure symbolic reasoning tasks to narrow the gap. However, the results indicate that FT-Llama2 can utilize similar template matching to respond to reasoning queries, but it falls short of generalizing to novel logic rules. These surprising observations question whether modern LLMs have mastered the inductive, deductive, and abductive reasoning abilities as in human intelligence, and motivate research on unveiling the magic existing within the black-box LLMs and evaluating and improving language models' reasoning abilities.
Author context
Most prolific author: 13 submissions (credibility 1.00).
Delta if applied: -0.2 percentile
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 36 comparisons
Ranked above opponent in 43% of matchups.
- ▼ lost to Beyond Size: How Gradients Shape Pruning D… ×4
- ▼ lost to Unified Language-Vision Pretraining in LLM… ×4
- ▼ lost to Empowering Active Learning for 3D Molecula… ×4
- ▼ lost to Gradient norm as a powerful proxy to out-o… ×4
- ▲ beat ACID: Abstractive, Content-Based IDs for D… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 36)