PapersWithELO
← ICLR 2024 leaderboard

It HAS to be Subjective: Human Annotator Simulation via Zero-shot Density Estimation

Wen Wu, Wenlin Chen, Chao Zhang, Phil Woodland

representation learninghuman annotator simulationnormalizing flowsspeech processingemotion recognitionmeta-learningzero-shot learningfairness
89.20100
Fused
band ≈ ±15 pct pts (from σ = 0.29)
93.00100
Mimo
band ≈ ±19 pct pts (from σ = 0.39)
83.50100
DeepSeek
band ≈ ±22 pct pts (from σ = 0.44)

OpenReview ground truth

Rejected

TL;DR — This paper studies human annotator simulation (HAS) for automatic data labelling and model evaluation, which incorporates the variability of human evaluations.

Abstract

Human annotator simulation (HAS) serves as a cost-effective substitute for human evaluation such as data annotation and system assessment. Human perception and behaviour during human evaluation exhibit inherent variability due to diverse cognitive processes and subjective interpretations, which should be taken into account in modelling to better mimic the way people perceive and interact with the world. This paper introduces a novel meta-learning framework that treats HAS as a zero-shot density estimation problem, which incorporates human variability and allows for the efficient generation of human-like annotations for unlabelled test inputs. Under this framework, we propose two new model classes, conditional integer flows and conditional softmax flows, to account for ordinal and categorical annotations, respectively. The proposed method is evaluated on three real-world human evaluation tasks and shows superior capability and efficiency to predict the aggregated behaviours of human annotators, match the distribution of human annotations, and simulate the inter-annotator disagreements.

Author context

Most prolific author: 4 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 38)