PapersWithELO
← ICLR 2024 leaderboard

Fairness Metric Impossibility: Investigating and Addressing Conflicts

Jake Robertson, Noor Awad, Thorsten L Schmidt, Frank Hutter

fairness, safety & privacyFairnessMulti-objective optimizationHyperparameter Optimization
13.20100
Fused
band ≈ ±15 pct pts (from σ = 0.29)
16.60100
Mimo
band ≈ ±22 pct pts (from σ = 0.43)
8.90100
DeepSeek
band ≈ ±19 pct pts (from σ = 0.39)

OpenReview ground truth

Rejected

Abstract

Fairness-aware ML (FairML) applications are often characterized by intricate social objectives and legal requirements, often encompassing multiple, potentially conflicting notions of fairness. Despite the well-known Impossibility Theorem of Fairness and vast theoretical research on the statistical and socio-technical trade-offs between fairness metrics, many FairML approaches still optimize for a single, user-defined fairness objective. However, this one-sided optimization can inadvertently lead to violations of other pertinent notions of fairness, resulting in adverse social consequences. In this exploratory and empirical study, we address the presence of fairness-metric conflicts by treating fairness metrics as conflicting objectives in a multi-objective (MO) sense. To efficiently explore multiple fairness-accuracy trade-offs and effectively balance conflicts between various fairness objectives, we introduce the ManyFairHPO framework, a novel many-objective (MaO) hyper-parameter optimization (HPO) approach. By enabling fairness practitioners to specify and explore complex and multiple fairness objectives, we open the door to further socio-technical research on effectively combining the complementary benefits of different notions of fairness.

Author context

Most prolific author: 6 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Battle history — 34 comparisons

Ranked above opponent in 33% of matchups.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 34)