PapersWithELO
← ICLR 2024 leaderboard

On the Paradox of Generalizable Logical Reasoning in Large Language Models

Xiaojuan Tang, Zilong Zheng, Jiaqi Li, Fanxu Meng, Song-Chun Zhu, Yitao Liang, Muhan Zhang

representation learningLarge Language ModelsSymbolic Reasoning
36.50100
Fused
band ≈ ±14 pct pts (from σ = 0.29)
40.20100
Mimo
band ≈ ±19 pct pts (from σ = 0.38)
36.40100
DeepSeek
band ≈ ±21 pct pts (from σ = 0.43)

OpenReview ground truth

Rejected

Abstract

The emergent few-shot reasoning capabilities of Large Language Models (LLMs) have excited the natural language and machine learning community over recent years. Despite the numerous successful applications, it remains an open question whether LLMs have generalizable logical reasoning abilities. In this work, we expose a surprising failure of generalization in logical reasoning tasks (deduction, induction, and abduction)---when semantics are decoupled from the language reasoning process (\ie, replacing semantic words with pure symbols), LLMs tend to perform much worse. We hypothesize that the learned \textit{semantics} of language tokens do the most heavy lifting during the reasoning process but fail to imitate the basic formal reasoning abilities of humans. Furthermore, we also attempt to fine-tune Llama-2 on pure symbolic reasoning tasks to narrow the gap. However, the results indicate that FT-Llama2 can utilize similar template matching to respond to reasoning queries, but it falls short of generalizing to novel logic rules. These surprising observations question whether modern LLMs have mastered the inductive, deductive, and abductive reasoning abilities as in human intelligence, and motivate research on unveiling the magic existing within the black-box LLMs and evaluating and improving language models' reasoning abilities.

Author context

Most prolific author: 13 submissions (credibility 1.00).

Delta if applied: -0.2 percentile

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Battle history — 36 comparisons

Ranked above opponent in 43% of matchups.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 36)