PapersWithELO
← ICLR 2024 leaderboard

Tell, Don't Show: Internalized Reasoning influences how LLMs generalize

Alexander Meinke, Owain Evans

generative modelsLLM generalizationAI safety
6.30100
Fused
band ≈ ±15 pct pts (from σ = 0.30)
12.90100
Mimo
band ≈ ±20 pct pts (from σ = 0.40)
3.60100
DeepSeek
band ≈ ±23 pct pts (from σ = 0.46)

OpenReview ground truth

Rejected

TL;DR — When training an LLM on examples that follow a pattern, but also giving the model declarative information stating the opposite of the pattern at train time, we investigate how the model ends up generalizing.

Abstract

In this paper we investigate to what extent language models' generalization behavior during a domain shift can be influenced by declarative knowledge contained in the training data. In order to study this we finetune language models to fit some distribution which has a ``natural'' generalization when the distribution shifts. We then test to what extent declarative statements in the training data - that if fully internalized would greatly affect the domain shift generalization - can indeed alter the model's behavior on unseen examples. While the effect is subtle, the declarative knowledge provided in the finetuning sets systematically changes the models' predictions in the way one would expect. Evidence for the strength of this effect growing with model size is mixed. We further show that the effect can not be explained by simple token matching behavior as it persists even when there is no overlap between the declarative descriptions and the models' test time generations.

Author context

Most prolific author: 4 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Battle history — 32 comparisons

Ranked above opponent in 41% of matchups.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 32)