PapersWithELO
← ICLR 2024 leaderboard

Exploring the State and Action Space in Reinforcement Learning with Infinite-Dimensional Confidence Balls

Yucong Lin, Yicheng Teng, Jingda Wu, Junwei Lu

reinforcement learningonline reinforcement learningreproducing kernel Hilbert spaceembedding learning
20.30100
Fused
band ≈ ±14 pct pts (from σ = 0.29)
17.30100
Mimo
band ≈ ±20 pct pts (from σ = 0.41)
34.00100
DeepSeek
band ≈ ±20 pct pts (from σ = 0.40)

OpenReview ground truth

Rejected

TL;DR — We propose a novel approach that leverages reproducing kernel Hilbert spaces (RKHSs) to tackle the curse of dimensionality problem on online reinforcement learning with continuous state and action spaces.

Abstract

Reinforcement Learning (RL) is a powerful tool for solving complex decision-making problems. However, existing RL approaches suffer from the curse of dimensionality when dealing with large or continuous state and action spaces. This paper introduces a non-parametric online RL algorithm called RKHS-RL that overcomes these challenges by utilizing reproducing kernels and the RKHS-embedding assumption. The proposed algorithm can handle both finite and infinite state and action spaces, as well as nonlinear relationships in transition probabilities. The RKHS-RL algorithm estimates the transition core using ridge regression and balances exploration and exploitation through infinite-dimensional confidence balls. The paper provides theoretical guarantees, demonstrating that RKHS-RL achieves a sublinear regret bound of $\tilde{\mathcal{O}}(H\sqrt{T})$, where $T$ denotes the time step of the algorithm and $H$ represents the horizon of the Markov Decision Process (MDP), making it an effective approach for RL problems.

Author context

Most prolific author: 3 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Battle history — 36 comparisons

Ranked above opponent in 44% of matchups.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 36)