Human Feedback is not Gold Standard
Tom Hosking, Phil Blunsom, Max Bartolo
OpenReview ground truth
TL;DR — We critically analyse the use of human feedback for evaluating and training Large Language Models, finding that human preference scores under-represent some crucial error types, and are biased by the assertiveness of the output.
Abstract
Human feedback has become the de facto standard for evaluating the performance of Large Language Models, and is increasingly being used as a training objective. However, it is not clear which properties of a generated output this single `preference' score captures. We hypothesise that preference scores are subjective and open to undesirable biases. We critically analyse the use of human feedback for both training and evaluation, to verify whether it fully captures a range of crucial error criteria. We find that while preference scores have fairly good coverage, they under-represent important aspects like factuality. We further hypothesise that both preference scores and error annotation may be affected by confounders, and leverage instruction-tuned models to generate outputs that vary along two possible confounding dimensions: assertiveness and complexity. We find that the assertiveness of an output skews the perceived rate of factuality errors, indicating that human annotations are not a fully reliable evaluation metric or training objective. Finally, we offer preliminary evidence that using human feedback as a training objective disproportionately increases the assertiveness of model outputs. We encourage future work to carefully consider whether preference scores are well aligned with the desired objective.
Author context
Most prolific author: 2 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 30 comparisons
Ranked above opponent in 55% of matchups.
- ▲ beat Compact Text-to-SDF via Latent Modeling ×4
- ▲ beat Learning-Retrieval-Revision For Large Lang… ×4
- ▲ beat Cross-domain Adaptation for Few-shot 3D Sh… ×4
- ▼ lost to In-Context Learning Learns Label Relations… ×4
- ▼ lost to Perceptual Group Tokenizer: Building Perce… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 30)