Focus on Primary: Differential Diverse Data Augmentation for Generalization in Visual Reinforcement Learning
Junhong Wu, Jie Liu, Xi Xiong, Daolong An, Shuai Lü
OpenReview ground truth
Abstract
In reinforcement learning, it is common for the agent to overfit the training environment, making generalization to unseen environments extremely challenging. Visual reinforcement learning that relies on observed images as input is particularly constrained by generalization and sample efficiency. To address these challenges, various data augmentation methods are consistently attempted to improve the generalization capability and reduce the training cost. However, the naive use of data augmentation can often lead to breakdowns in learning. In this paper, we propose two novel approaches: Diverse Data Augmentation (DDA) and Differential Diverse Data Augmentation (D3A). Leveraging a pre-trained encoder-decoder model, we segment primary pixels to avoid inappropriate data augmentation affecting critical information. DDA improves the generalization capability of the agent in complex environments through consistency of encoding. D3A uses proper data augmentation for primary pixels to further improve generalization while satisfying semantic-invariant state transformation. We extensively evaluate our methods on a series of generalization tasks of DeepMind Control Suite. The results demonstrate that our methods significantly improve the generalization performance of the agent in unseen environments, and enable the selection of more diverse data augmentations to improve the sample efficiency of off-policy algorithms.
Author context
Most prolific author: 2 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 36 comparisons
Ranked above opponent in 45% of matchups.
- ▼ lost to HIPODE: Enhancing Offline Reinforcement Le… ×10
- ▼ lost to Interpreting Categorical Distributional Re… ×4
- ▼ lost to Maximum Entropy On-Policy Actor-Critic via… ×4
- ▼ lost to Towards Assessing and Benchmarking Risk-Re… ×4
- ▲ beat Unraveling the Enigma of Double Descent: A… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 36)