← ICLR 2024 leaderboard

Query and Response Augmentation Cannot Help Out-of-domain Math Reasoning Generalization

Chengpeng Li, Zheng Yuan, Hongyi Yuan, Guanting Dong, Keming Lu, Jiancan Wu, Chuanqi Tan, Xiang Wang, Chang Zhou

representation learningLarge Language ModelMath ReasoningData augmentationScaling relationshipGeneralizability
28.10100
Fused
band ≈ ±13 pct pts (from σ = 0.27)
34.80100
Mimo
band ≈ ±19 pct pts (from σ = 0.38)
22.40100
DeepSeek
band ≈ ±19 pct pts (from σ = 0.37)

OpenReview ground truth

Rejected

TL;DR — This paper analyzes the scaling relationship and generalization of data augmentation in mathematical reasoning with large language models.

Abstract

In math reasoning with large language models (LLMs), fine-tuning data augmentation by query evolution and diverse reasoning paths is empirically verified effective, profoundly narrowing the gap between open-sourced LLMs and cutting-edge proprietary LLMs. In this paper, we conduct an investigation for such data augmentation in math reasoning and are intended to answer: (1) What strategies of data augmentation are more effective; (2) What is the scaling relationship between the amount of augmented data and model performance; and (3) Can data augmentation incentivize generalization to out-of-domain mathematical reasoning tasks? To this end, we create a new dataset, AugGSM8K, by complicating and diversifying the queries from GSM8K and sampling multiple reasoning paths. We obtained a series of LLMs called MuggleMath by fine-tuning on subsets of AugGSM8K. MuggleMath substantially achieves new state-of-the-art on GSM8K (from 54\% to 68.4\% at the scale of 7B, and from 63.9\% to 74.0\% at the scale of 13B). A log-linear relationship is presented between MuggleMath’s performance and the amount of augmented data. We also find that MuggleMath is weak in out-of-domain math reasoning generalization to MATH. This is attributed to the differences in query distribution between AugGSM8K and MATH which suggest that augmentation on a single benchmark could not help with overall math reasoning performance.

Author context

Most prolific author: 9 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Battle history — 44 comparisons

Ranked above opponent in 51% of matchups.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 44)