PapersWithELO
← ICLR 2024 leaderboard

HIPODE: Enhancing Offline Reinforcement Learning with High-Quality Synthetic Data from a Policy-Decoupled Approach

Shixi Lian, Jianye HAO, Yi Ma, Jinyi Liu, YAN ZHENG

reinforcement learningData AugmentationOffline Reinforcement Learning
21.00100
Fused
band ≈ ±14 pct pts (from σ = 0.28)
23.90100
Mimo
band ≈ ±20 pct pts (from σ = 0.40)
20.00100
DeepSeek
band ≈ ±19 pct pts (from σ = 0.39)

OpenReview ground truth

Rejected

TL;DR — We propose a novel data augmentation algorithm HIPODE for offline RL to generate high-return synthetic data, which is beneficial for different downstream policies.

Abstract

Offline reinforcement learning (Offline RL) has gained attention as a means of training reinforcement learning models using pre-collected static data. To address the issue of limited data and improve downstream Offline RL performance, recent efforts have focused on broadening dataset coverage through data augmentation techniques. However, most of these methods are tied to a specific policy (policy-dependent), restricting the generated data to supporting only a specific downstream Offline RL policy. Moreover, the quality of synthetic data is often not well-controlled, which limits the potential for further improving the downstream policy. To tackle these issues, we propose \textbf{HI}gh-return \textbf{PO}licy-\textbf{DE}coupled~(HIPODE), a novel data augmentation method for Offline RL. On the one hand, HIPODE generates high-return synthetic data by selecting states near the dataset distribution with potentially high value among candidate states using the negative sampling technique. On the other hand, HIPODE is policy-decoupled, thus can be used as a common plug-in method to support diverse downstream Offline RL processes. We conduct experiments on the widely studied TD3BC, CQL and IQL algorithms, and the results show that HIPODE outperforms or has competitive results to the state-of-the-art policy-decoupled data augmentation method and most prevalent model-based Offline RL methods on D4RL benchmarks.

Author context

Most prolific author: 10 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Battle history — 36 comparisons

Ranked above opponent in 44% of matchups.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 36)