Semi-supervised batch learning from logged data
Gholamali Aminian, Armin Behnamnia, Roberto Vega, Laura Toni, Chengchun Shi, Hamid R. Rabiee, Omar Rivasplata, Miguel R. D. Rodrigues
OpenReview ground truth
TL;DR — We propose a theoretical-inspired semi-supervised batch learning from logged data with known-reward and missing-reward samples.
Abstract
Offline policy learning methods are intended to learn a policy from logged data, which includes context, action, and reward for each sample point. In this work we build on the counterfactual risk minimization framework, which also assumes access to propensity scores. We propose learning methods for problems where rewards of some samples are missing, so there are samples with rewards and samples missing rewards in the logged data. We refer to this type of learning as semi-supervised batch learning from logged data, which arises in a wide range of application domains. We derive new upper bound for the true risk under inverse propensity score estimation to better address this kind of learning problem. Using this bound, we propose a regularized semi-supervised batch learning method with logged data where the regularization term is reward-independent and, as a result, can be evaluated using the logged missing-reward data. Consequently, even though reward feedback is only present for some samples, a parameterized policy can be learned by leveraging the missing-reward samples. The results of experiments derived from benchmark datasets indicate that these algorithms achieve policies with better performance in comparison with logging policies.
Author context
Most prolific author: 3 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 42 comparisons
Ranked above opponent in 41% of matchups.
- ▼ lost to Physics-Regulated Deep Reinforcement Learn… ×6
- ▲ beat HiLoRL: A Hierarchical Logical Model for L… ×6
- ▼ lost to The Curse of Diversity in Ensemble-Based E… ×4
- ▼ lost to Identifying Latent State Transition Proces… ×4
- ▼ lost to On Stationary Point Convergence of PPO-Cli… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 42)