A Symbolic Framework for Evaluating Mathematical Reasoning with Transformers
Jordan Meadows, Marco Valentino, Damien Teney, Andre Freitas
OpenReview ground truth
TL;DR — A symbolic data generation and perturbation framework is proposed and employed to determine differences in mathematical generalisation capabilities between fine-tuned BERT and few-shot GPT models, in a number of sequence classification tasks.
Abstract
This paper proposes a methodology for generating synthetic mathematical derivations via a computer algebra system to evaluate the generalisability of Transformers in symbolic and quantitative reasoning problems, and provides a general framework for building large-scale and high-quality benchmarks in the mathematical domain. In the context of classification tasks involving multi-step annotated derivations (spanning 18 mathematical operators), we leverage the framework to compare the mathematical capabilities of GPT-4, GPT-3.5, and a canon of fine-tuned BERT models, exploring the relationship between specific operators and generalisation failure. Surprisingly, the average in-distribution performance of BERT models surpasses GPT-3.5, and rivals GPT-4, yet simple data perturbations reduce BERT scores by up to 80 F1 points. The results suggest that the in-distribution performance and generalisability of smaller open-source models may potentially rival GPT in narrow mathematical domains by incorporating appropriately structured discourse-level relations during training, and highlight a shared weakness between BERT and GPT involving a relative inability to decode dependency relations involving indirect references to mathematical entities. We release the data generation framework along with all the resulting datasets and fine-tuned models\footnote{\url{https://github.com/anonymous/TBA}}.
Author context
Most prolific author: 4 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 38 comparisons
Ranked above opponent in 43% of matchups.
- ▼ lost to CompA: Addressing the Gap in Compositional… ×6
- ▲ beat OpenReviewer: Mitigating Challenges in LLM… ×4
- ▼ lost to SOTOPIA: Interactive Evaluation for Social… ×4
- ▲ beat SKILL-MIX: a Flexible and Expandable Famil… ×4
- ▼ lost to What Makes ImageNet Look Unlike LAION ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 38)