TIGERScore: Building Explainable Metric for All Text Generation Task
Dongfu Jiang, Yishan Li, Ge Zhang, Wenhao Huang, Bill Yuchen Lin, Wenhu Chen
OpenReview ground truth
TL;DR — We present TIGERScore, an instruction-guided evaluation scorer aiming to evaluate any text generation for various tasks in an explainable format
Abstract
We present TIGERScore, a \textbf{T}rained metric that follow \textbf{I}nstruction \textbf{G}uidance to perform \textbf{E}xplainable, and \textbf{R}eference-free evaluation over a wide spectrum of text generation tasks. Different from other automatic evaluation methods that only provide arcane scores, TIGERScore is guided by the natural language instruction to provide error analysis to pinpoint the mistakes in the generated text. Our metric is based on LLaMA, trained on our meticulously curated instruction-tuning dataset MetricInstruct that covers 6 text generation tasks and 23 text generation datasets. The dataset consists of 48K quadruple in the form of (instruction, input, system output $\rightarrow$ error analysis). We collected the `system outputs' through diverse channels to cover different types of errors. To quantitatively assess our metric, we evaluate its correlation with human ratings on 5 held-in datasets, 2 held-out datasets and show that TIGERScore can achieve the highest overall Spearman's correlation with human ratings across these datasets and outperforms other metrics significantly. As a reference-free metric, its correlation can even surpass the best existing reference-based metrics. To further qualitatively assess the rationale generated by our metric, we conduct human evaluation on the generated explanations and found that the explanations are 70.8\% accurate. Through these experimental results, we believe TIGERScore demonstrate the possibilities to build a universal explainable metrics to evaluate any text generation task.
Author context
Most prolific author: 6 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 34 comparisons
Ranked above opponent in 47% of matchups.
- ▼ lost to Reinforced UI Instruction Grounding: Towar… ×6
- ▼ lost to PromptAgent: Strategic Planning with Langu… ×4
- ▲ beat Memorization for Good: Encryption with Aut… ×4
- ▲ beat ProtoReg: Prioritizing Discriminative Info… ×4
- ▼ lost to Learn from the Past: A Proxy based Adversa… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 34)