PapersWithELO
← ICLR 2024 leaderboard

Variational Language Concepts for Interpreting Pretrained Language Models

Hengyi Wang, Zhiqing Hong, Desheng Zhang, Hao Wang

probabilistic methodsConceptual InterpretationInterpretability and ExplainabilityGenerative ModelsProbabilistic Graphical ModelsPretrained Lauguage Models
50.80100
Fused
band ≈ ±14 pct pts (from σ = 0.29)
62.50100
Mimo
band ≈ ±21 pct pts (from σ = 0.41)
37.90100
DeepSeek
band ≈ ±20 pct pts (from σ = 0.40)

OpenReview ground truth

Rejected

TL;DR — We propose a hierarchical variational Bayesian framework to provide conceptual interpretations of pretrained language models.

Abstract

Pretrained Language Models (PLMs) such as BERT and its variants have achieved remarkable success in natural language processing. To date, the interpretability of PLMs has primarily relied on the attention weights in their self-attention layers. However, these attention weights only provide word-level interpretations, failing to capture higher-level structures, and are therefore lacking in readability and intuitiveness. To address this challenge, we first provide a formal definition of ``conceptual interpretation`` and then propose a variational Bayesian framework, dubbed VAriational LANguage ConcEpt (VALANCE), to go beyond word-level interpretations and provide concept-level interpretations. Our theoretical analysis shows that our VALANCE finds the optimal language concepts to interpret PLM predictions. Empirical results on several real-world datasets show that our method can successfully provide conceptual interpretation for PLMs.

Author context

Most prolific author: 7 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 36)