LMExplainer: A Knowledge-Enhanced Explainer for Language Models
Zichen Chen, Jianda Chen, Chen YuanYuan, Han Yu, Ambuj Singh, Misha Sra
OpenReview ground truth
Abstract
Language models (LMs), such as GPT-4, are powerful tools for natural language processing, capable of handling diverse tasks from text generation to question answering. However, their decision process lack transparency due to the complex, multi-layered, and nonlinear model structures involving millions of parameters. This hinders user trust on LMs, especially in safety-critical applications. Due to the opaque nature of LMs, a promising approach for explaining how they work is by generating explanations on a more transparent surrogate (e.g., a knowledge graph (KG)). Such works mostly exploit attention weights to provide explanations for LM recommendations. However, pure attention-based explanations lack scalability to keep up with the growing complexity of LMs. To bridge this important gap, we propose LMExplainer, a knowledge-enhanced explainer for LMs capable of providing human-understandable explanations. It is designed to efficiently locate the most relevant knowledge within a large-scale KG via the graph attention neural network (GAT) to extract key decision signals reflecting how a given LM works. Extensive experiments comparing LMExplainer against seven state-of-the-art baselines show that it outperforms existing LM+KG methods on the CommonsenseQA and OpenBookQA datasets. We compare the explanation generated by LMExplainer with other algorithm-generated explanations as well as human-annotated explanations. The results show that LMExplainer generates more comprehensive and clearer explanations.
Author context
Most prolific author: 5 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 32 comparisons
Ranked above opponent in 36% of matchups.
- ▲ beat Mitigating Accumulated Distribution Diverg… ×6
- ▼ lost to Everybody Needs a Little HELP: Explaining … ×4
- ▼ lost to SPADE: Sparsity-Guided Debugging for Deep … ×4
- ▼ lost to Function Vectors in Large Language Models ×4
- ▼ lost to PRIME: Prioritizing Interpretability in Fa… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 32)