PapersWithELO
← ICLR 2024 leaderboard

DEXR: A Unified Approach Towards Environment Agnostic Exploration

Yiran Wang, Yunfan Li, Sanae Amani, Lin Yang

reinforcement learningReinforcement LearningExplorationIntrinsic Rewards
23.50100
Fused
band ≈ ±16 pct pts (from σ = 0.31)
43.50100
Mimo
band ≈ ±22 pct pts (from σ = 0.43)
13.80100
DeepSeek
band ≈ ±22 pct pts (from σ = 0.44)

OpenReview ground truth

Rejected

TL;DR — We propose a novel framework DEXR for enhancing exploration algorithms across different types of environments, by preventing over-exploratiton and optimization instability.

Abstract

The exploration-exploitation dilemma poses pivotal challenges in reinforcement learning (RL). While recent advances in curiosity-driven techniques have demonstrated capabilities in sparse reward scenarios, they necessitate extensive hyperparameter tuning on different types of environments and often fall short in dense reward settings. In response to these challenges, we introduce the novel \textbf{D}elayed \textbf{EX}ploration \textbf{R}einforcement Learning (DEXR) framework. DEXR adeptly curbs over-exploration and optimization instabilities issues of curiosity-driven methods, and can efficiently adapt to both dense and sparse reward environments with minimal hyperparameter tuning. This is facilitated by an auxiliary exploitation-only policy that streamlines data collection, guiding the exploration policy towards high-value regions and minimizing unnecessary exploration. Additionally, this exploration policy yields diverse, in-distribution data, and bolsters training robustness with neural network structures. We verify the efficacy of DEXR with both theoretical validations and comprehensive empirical evaluations, demonstrating its superiority in a broad range of environments.

Author context

Most prolific author: 1 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Battle history — 32 comparisons

Ranked above opponent in 46% of matchups.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 32)