PapersWithELO
← ICLR 2024 leaderboard

DeCCaF: Deferral Under Cost and Capacity Constraints Framework

Jean Vieira Alves, Diogo Leitão, Sérgio Jesus, Marco O. P. Sampaio, Pedro Saleiro, Mario A. T. Figueiredo, Pedro Bizarro

general MLLearning to DeferHuman-AI TeamingHuman-AI Collaboration
32.50100
Fused
band ≈ ±14 pct pts (from σ = 0.29)
45.50100
Mimo
band ≈ ±20 pct pts (from σ = 0.40)
20.60100
DeepSeek
band ≈ ±21 pct pts (from σ = 0.41)

OpenReview ground truth

Rejected

Abstract

The \textit{learning to defer} (L2D) framework aims to improve human-AI collaboration systems by deferring decisions to humans when they are more likely to make the correct judgment than a ML classifier. Existing research in L2D overlooks key aspects of real-world systems that impede its practical adoption, such as: i) neglecting cost-sensitive scenarios; ii) requiring concurrent human predictions for every instance of the dataset in training and iii) not dealing with human capacity constraints. To address these issues, we propose the \textit{deferral under cost and capacity constraint framework} (DeCCaF). A novel L2D approach: DeCCaF employs supervised learning to model the probability of human error with less restrictive data requirements (only one expert prediction per instance), and uses constraint programming to globally minimize error cost subject to capacity constraints. We employ DeCCaF in a cost-sensitive fraud detection setting with a team of 50 synthetic fraud analysts, subject to a wide array of realistic human work capacity constraints, showing that DeCCaF significantly outperforms L2D baselines, reducing average misclassification costs by 9 \%. Our code and testbed are available at https://anonymous.4open.science/r/deccaf-1245/

Author context

Most prolific author: 1 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Battle history — 34 comparisons

Ranked above opponent in 45% of matchups.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 34)