PapersWithELO
← ICLR 2024 leaderboard

Learning to Reject Meets Long-tail Learning

Harikrishna Narasimhan, Aditya Krishna Menon, Wittawat Jitkrittum, Neha Gupta, Sanjiv Kumar

learning theoryLearning to rejectbalanced errorevaluation metricsselective classificationplug-in approachlong-tail learningclass imbalancenon-decomposable metrics
95.80100
Fused
band ≈ ±15 pct pts (from σ = 0.31)
97.10100
Mimo
band ≈ ±23 pct pts (from σ = 0.46)
90.40100
DeepSeek
band ≈ ±20 pct pts (from σ = 0.41)

OpenReview ground truth

Accepted

TL;DR — Theoretically-grounded approach for learning to reject with general evaluation metrics including the balanced and worst-group errors

Abstract

Learning to reject (L2R) is a classical problem where one seeks a classifier capable of abstaining on low-confidence samples. Most prior work on L2R has focused on minimizing the standard misclassification error. However, in many real-world applications, the label distribution is highly imbalanced, necessitating alternate evaluation metrics such as the balanced error or the worst-group error that enforce equitable performance across both the head and tail classes. In this paper, we establish that traditional L2R methods can be grossly sub-optimal for such metrics, and show that this is due to an intricate dependence in the objective between the label costs and the rejector. We then derive the form of the Bayes-optimal classifier and rejector for the balanced error, propose a novel plug-in approach to mimic this solution, and extend our results to general evaluation metrics. Through experiments on benchmark image classification tasks, we show that our approach yields better trade-offs in both the balanced and worst-group error compared to L2R baselines.

Author context

Most prolific author: 10 submissions (credibility 0.97).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 36)