PapersWithELO
← ICLR 2024 leaderboard

Multiobjective Stochastic Linear Bandits under Lexicographic Ordering

Bo Xue, Xi Lin, Xiaoyuan Zhang, Qingfu Zhang

reinforcement learningmultiobjectivebanditslexicographic ordering
88.10100
Fused
band ≈ ±15 pct pts (from σ = 0.29)
85.20100
Mimo
band ≈ ±20 pct pts (from σ = 0.41)
89.60100
DeepSeek
band ≈ ±21 pct pts (from σ = 0.42)

OpenReview ground truth

Rejected

TL;DR — We develop an almost optimal algorithm for the multiobjective stochastic linear bandits under lexicographic ordering.

Abstract

This paper studies the multiobjective stochastic linear bandit (MOSLB) model under lexicographic ordering, where the agent aims to simultaneously maximize $m$ objectives in a hierarchical manner. This model has various real-world scenarios, including water resource planning and radiation treatment for cancer patients. However, there is no effort on the general MOSLB model except a special case called multiobjective multi-armed bandits. Previous literature provided a suboptimal algorithm for this special case, which enjoys a regret bound of $\widetilde{O}(T^{2/3})$ under a priority-based regret measure. In this paper, we propose an algorithm achieving the almost optimal regret bound $\widetilde{O}(d\sqrt{T})$ for the MOSLB model, and its metric is the general regret. Here, $d$ is the dimension of arm vector and $T$ is the time horizon. The major novelties of our algorithm include a new arm filter and a multiple trade-off approach for exploration and exploitation. Experiments confirm the merits of our algorithms and provide compelling evidence to support our analysis.

Author context

Most prolific author: 5 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 36)