Tell, Don't Show: Internalized Reasoning influences how LLMs generalize
Alexander Meinke, Owain Evans
OpenReview ground truth
TL;DR — When training an LLM on examples that follow a pattern, but also giving the model declarative information stating the opposite of the pattern at train time, we investigate how the model ends up generalizing.
Abstract
In this paper we investigate to what extent language models' generalization behavior during a domain shift can be influenced by declarative knowledge contained in the training data. In order to study this we finetune language models to fit some distribution which has a ``natural'' generalization when the distribution shifts. We then test to what extent declarative statements in the training data - that if fully internalized would greatly affect the domain shift generalization - can indeed alter the model's behavior on unseen examples. While the effect is subtle, the declarative knowledge provided in the finetuning sets systematically changes the models' predictions in the way one would expect. Evidence for the strength of this effect growing with model size is mixed. We further show that the effect can not be explained by simple token matching behavior as it persists even when there is no overlap between the declarative descriptions and the models' test time generations.
Author context
Most prolific author: 4 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 32 comparisons
Ranked above opponent in 41% of matchups.
- ▼ lost to Learning Energy-Based Models by Cooperativ… ×4
- ▼ lost to On the Evaluation of Generative Models in … ×4
- ▼ lost to Attribute Based Interpretable Evaluation M… ×4
- ▼ lost to Potential Based Diffusion Motion Planning ×4
- ▼ lost to TeLLMe what you see: Using LLMs to Explain… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 32)