← ICLR 2024 leaderboard

Knowledge Manipulation in Language Models (Part B)

Zeyuan Allen-Zhu, Yuanzhi Li

interpretability & vizInterpretabilityTransformersLanguage ModelsLinear ProbingInner WorkingFactual Knowledge
46.10100
Fused
band ≈ ±15 pct pts (from σ = 0.31)
56.10100
Mimo
band ≈ ±21 pct pts (from σ = 0.41)
27.90100
DeepSeek
band ≈ ±23 pct pts (from σ = 0.45)

OpenReview ground truth

Rejected

TL;DR — Why do LLMs need Chain of Thoughts even for basic questions (e.g. was Biden born on an even day)? We show that LLMs cannot efficiently manipulate knowledge even if such knowledge is 100% extractable; plus, inverse knowledge search is just impossible.

Abstract

Language models can store vast amounts of factual knowledge, but their ability to use this knowledge for logical reasoning remains questionable. This paper explores a language model's ability to manipulate its stored knowledge during inference. We focus on four manipulation types: *retrieval* (e.g., "What is person A's attribute X"), *classification* (e.g., "Is A's attribute X even or odd?"), *comparison* (e.g., "Is A greater than B in attribute X?") and *inverse search* (e.g., "Which person's attribute X equals T?") We observe that pre-trained language models like GPT2/3/4 excel in knowledge retrieval but struggle with simple classification or comparison tasks unless Chain of Thoughts (CoTs) are employed during both training and inference. They also perform poorly in inverse knowledge search, irrespective of the prompts. Our primary contribution is a synthetic dataset for a *controlled experiment* that confirms these inherent weaknesses: a language model cannot *efficiently* manipulate knowledge from pre-training data, even when such knowledge is perfectly stored and fully extractable in the models, and despite adequate instruct fine-tuning.

Author context

Most prolific author: 12 submissions (credibility 0.37).

Delta if applied: -0.6 percentile

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 34)