Exploring the State and Action Space in Reinforcement Learning with Infinite-Dimensional Confidence Balls
Yucong Lin, Yicheng Teng, Jingda Wu, Junwei Lu
OpenReview ground truth
TL;DR — We propose a novel approach that leverages reproducing kernel Hilbert spaces (RKHSs) to tackle the curse of dimensionality problem on online reinforcement learning with continuous state and action spaces.
Abstract
Reinforcement Learning (RL) is a powerful tool for solving complex decision-making problems. However, existing RL approaches suffer from the curse of dimensionality when dealing with large or continuous state and action spaces. This paper introduces a non-parametric online RL algorithm called RKHS-RL that overcomes these challenges by utilizing reproducing kernels and the RKHS-embedding assumption. The proposed algorithm can handle both finite and infinite state and action spaces, as well as nonlinear relationships in transition probabilities. The RKHS-RL algorithm estimates the transition core using ridge regression and balances exploration and exploitation through infinite-dimensional confidence balls. The paper provides theoretical guarantees, demonstrating that RKHS-RL achieves a sublinear regret bound of $\tilde{\mathcal{O}}(H\sqrt{T})$, where $T$ denotes the time step of the algorithm and $H$ represents the horizon of the Markov Decision Process (MDP), making it an effective approach for RL problems.
Author context
Most prolific author: 3 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 36 comparisons
Ranked above opponent in 44% of matchups.
- ▲ beat Generalized Convergence Analysis of Tsetli… ×6
- ▼ lost to The Update-Equivalence Framework for Decis… ×4
- ▼ lost to DiffCPS: Diffusion Model based Constrained… ×4
- ▼ lost to Continual Offline Reinforcement Learning v… ×4
- ▼ lost to CLIP as Multi-Task Multi-Kernel Learning ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 36)