Maximizing Benefits under Harm Constraints: A Generalized Linear Contextual Bandit Approach
Yingchao Zhou, Jia Liu, Zhengyuan Zhu
OpenReview ground truth
Abstract
In many contextual sequential decision-making scenarios, such as dose-finding clinical trials for new drugs or personalized news article recommendation systems in social media, each action can simultaneously carry both benefits and potential harm. This could manifest as efficacy versus side effects in clinical trials, or increased user engagement versus the risk of radicalization and psychological distress in news recommendation. These multifaceted situations can be modeled using the multi-armed bandit (MAB) framework. Given the intricate balance of positive and negative outcomes in these contexts, there is a compelling need to develop methods which can maximize benefits while limiting harm within the MAB framework. This paper aims to address this gap. The primary contributions of this paper are two-fold: (i) We propose a novel contextual MAB model with the objective of optimizing reward potential while maintaining certain harm constraints. In this model both rewards and harm are governed by a generalized linear model with coefficients that vary based on the contextual variables. This flexibility allows the model to be broadly applicable for a wide range of scenarios. (ii) Building on our proposed generalized linear contextual MAB model, we develop an $\epsilon$-greedy-based policy. This policy is designed to strike an effective balance between the dual objectives of exploration-exploitation to achieve the desired trade-off between benefit and harm. We demonstrate that this policy achieves a sublinear $\mathcal{O}(\sqrt{T\log T})$ regret. Extensive experimental results are presented to support our theoretical analyses and validate the effectiveness of our proposed model and policy.
Author context
Most prolific author: 7 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 38 comparisons
Ranked above opponent in 55% of matchups.
- ▲ beat LatentCBF: A Control Barrier Function in L… ×6
- ▲ beat Learning interpretable control inputs and … ×6
- ▼ lost to Privileged Sensing Scaffolds Reinforcement… ×4
- ▼ lost to Non-Asymptotic Analysis for Single-Loop (N… ×4
- ▲ beat Closing the gap on tabular data with Fouri… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 38)