PapersWithELO
← ICLR 2024 leaderboard

Semi-supervised batch learning from logged data

Gholamali Aminian, Armin Behnamnia, Roberto Vega, Laura Toni, Chengchun Shi, Hamid R. Rabiee, Omar Rivasplata, Miguel R. D. Rodrigues

reinforcement learningSemi-supervised batch learningoff-policy learningIPS estimatorlearning bounds
23.90100
Fused
band ≈ ±13 pct pts (from σ = 0.27)
19.30100
Mimo
band ≈ ±19 pct pts (from σ = 0.38)
22.50100
DeepSeek
band ≈ ±19 pct pts (from σ = 0.38)

OpenReview ground truth

Rejected

TL;DR — We propose a theoretical-inspired semi-supervised batch learning from logged data with known-reward and missing-reward samples.

Abstract

Offline policy learning methods are intended to learn a policy from logged data, which includes context, action, and reward for each sample point. In this work we build on the counterfactual risk minimization framework, which also assumes access to propensity scores. We propose learning methods for problems where rewards of some samples are missing, so there are samples with rewards and samples missing rewards in the logged data. We refer to this type of learning as semi-supervised batch learning from logged data, which arises in a wide range of application domains. We derive new upper bound for the true risk under inverse propensity score estimation to better address this kind of learning problem. Using this bound, we propose a regularized semi-supervised batch learning method with logged data where the regularization term is reward-independent and, as a result, can be evaluated using the logged missing-reward data. Consequently, even though reward feedback is only present for some samples, a parameterized policy can be learned by leveraging the missing-reward samples. The results of experiments derived from benchmark datasets indicate that these algorithms achieve policies with better performance in comparison with logging policies.

Author context

Most prolific author: 3 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Battle history — 42 comparisons

Ranked above opponent in 41% of matchups.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 42)