PapersWithELO
← ICLR 2024 leaderboard

Measuring Information in Text Explanations

Zining Zhu, Frank Rudzicz

representation learningexplanationinformationcommunication channel
20.90100
Fused
band ≈ ±14 pct pts (from σ = 0.27)
27.00100
Mimo
band ≈ ±20 pct pts (from σ = 0.40)
19.50100
DeepSeek
band ≈ ±19 pct pts (from σ = 0.37)

OpenReview ground truth

Rejected

Abstract

Text-based explanation is a particularly promising approach in explainable AI, but the evaluation of text explanations is method-dependent. We argue that placing the explanations on an information-theoretic framework could unify the evaluations of two popular text explanation methods: rationale and natural language explanations (NLE). This framework considers the post-hoc text pipeline as a series of communication channels, which we refer to as ``explanation channels''. We quantify the information flow through these channels, thereby facilitating the assessment of explanation characteristics. We set up tools for quantifying two information scores: relevance and informativeness. We illustrate what our proposed information scores measure by comparing them against some traditional evaluation metrics. Our information-theoretic scores reveal some unique observations about the underlying mechanisms of two representative text explanations. For example, the NLEs trade-off slightly between transmitting the input-related information and the target-related information, whereas the rationales do not exhibit such a trade-off mechanism. Our work contributes to the ongoing efforts in establishing rigorous and standardized evaluation criteria in the rapidly evolving field of explainable AI.

Author context

Most prolific author: 2 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Battle history — 38 comparisons

Ranked above opponent in 44% of matchups.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 38)