HIPODE: Enhancing Offline Reinforcement Learning with High-Quality Synthetic Data from a Policy-Decoupled Approach
Shixi Lian, Jianye HAO, Yi Ma, Jinyi Liu, YAN ZHENG
OpenReview ground truth
TL;DR — We propose a novel data augmentation algorithm HIPODE for offline RL to generate high-return synthetic data, which is beneficial for different downstream policies.
Abstract
Offline reinforcement learning (Offline RL) has gained attention as a means of training reinforcement learning models using pre-collected static data. To address the issue of limited data and improve downstream Offline RL performance, recent efforts have focused on broadening dataset coverage through data augmentation techniques. However, most of these methods are tied to a specific policy (policy-dependent), restricting the generated data to supporting only a specific downstream Offline RL policy. Moreover, the quality of synthetic data is often not well-controlled, which limits the potential for further improving the downstream policy. To tackle these issues, we propose \textbf{HI}gh-return \textbf{PO}licy-\textbf{DE}coupled~(HIPODE), a novel data augmentation method for Offline RL. On the one hand, HIPODE generates high-return synthetic data by selecting states near the dataset distribution with potentially high value among candidate states using the negative sampling technique. On the other hand, HIPODE is policy-decoupled, thus can be used as a common plug-in method to support diverse downstream Offline RL processes. We conduct experiments on the widely studied TD3BC, CQL and IQL algorithms, and the results show that HIPODE outperforms or has competitive results to the state-of-the-art policy-decoupled data augmentation method and most prevalent model-based Offline RL methods on D4RL benchmarks.
Author context
Most prolific author: 10 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 36 comparisons
Ranked above opponent in 44% of matchups.
- ▲ beat Focus on Primary: Differential Diverse Dat… ×10
- ▼ lost to Towards Assessing and Benchmarking Risk-Re… ×6
- ▼ lost to On Trajectory Augmentations for Off-Policy… ×6
- ▼ lost to Multi-Resolution Learning with DeepONets a… ×6
- ▼ lost to Interpreting Categorical Distributional Re… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 36)