DEXR: A Unified Approach Towards Environment Agnostic Exploration
Yiran Wang, Yunfan Li, Sanae Amani, Lin Yang
OpenReview ground truth
TL;DR — We propose a novel framework DEXR for enhancing exploration algorithms across different types of environments, by preventing over-exploratiton and optimization instability.
Abstract
The exploration-exploitation dilemma poses pivotal challenges in reinforcement learning (RL). While recent advances in curiosity-driven techniques have demonstrated capabilities in sparse reward scenarios, they necessitate extensive hyperparameter tuning on different types of environments and often fall short in dense reward settings. In response to these challenges, we introduce the novel \textbf{D}elayed \textbf{EX}ploration \textbf{R}einforcement Learning (DEXR) framework. DEXR adeptly curbs over-exploration and optimization instabilities issues of curiosity-driven methods, and can efficiently adapt to both dense and sparse reward environments with minimal hyperparameter tuning. This is facilitated by an auxiliary exploitation-only policy that streamlines data collection, guiding the exploration policy towards high-value regions and minimizing unnecessary exploration. Additionally, this exploration policy yields diverse, in-distribution data, and bolsters training robustness with neural network structures. We verify the efficacy of DEXR with both theoretical validations and comprehensive empirical evaluations, demonstrating its superiority in a broad range of environments.
Author context
Most prolific author: 1 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 32 comparisons
Ranked above opponent in 46% of matchups.
- ▼ lost to Provable and Practical: Efficient Explorat… ×4
- ▼ lost to Sample Efficient Reinforcement Learning fr… ×4
- ▼ lost to Compound Returns Reduce Variance in Reinfo… ×4
- ▼ lost to Multiobjective Stochastic Linear Bandits u… ×4
- ▲ beat Information based explanation methods for … ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 32)