PapersWithELO
← ICLR 2024 leaderboard

Efficient Action Robust Reinforcement Learning with Probabilistic Policy Execution Uncertainty

Guanlin Liu, Zhihan Zhou, Han Liu, Lifeng Lai

reinforcement learningRobust Reinforcement LearningSample Complexity
95.10100
Fused
band ≈ ±17 pct pts (from σ = 0.33)
93.40100
Mimo
band ≈ ±23 pct pts (from σ = 0.46)
92.70100
DeepSeek
band ≈ ±24 pct pts (from σ = 0.48)

OpenReview ground truth

Rejected

TL;DR — Introduce an efficient action robust RL algorithm that achieves minimax optimal regret and sample complexity.

Abstract

Robust reinforcement learning (RL) aims to find a policy that optimizes the worst-case performance in the face of uncertainties. In this paper, we focus on action robust RL with the probabilistic policy execution uncertainty, in which, instead of always carrying out the action specified by the policy, the agent will take the action specified by the policy with probability $1-\rho$ and an alternative adversarial action with probability $\rho$. We establish the existence of an optimal policy on the action robust MDPs with probabilistic policy execution uncertainty and provide the action robust Bellman optimality equation for its solution. Furthermore, we develop Action Robust Reinforcement Learning with Certificates (ARRLC) algorithm that achieves minimax optimal regret and sample complexity. Furthermore, we conduct numerical experiments to validate our approach's robustness, demonstrating that ARRLC outperforms non-robust RL algorithms and converges faster than the robust TD algorithm in the presence of action perturbations.

Author context

Most prolific author: 4 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 30)