PapersWithELO
← ICLR 2024 leaderboard

Continual Offline Reinforcement Learning via Diffusion-based Dual Generative Replay

Jinmei Liu, Wenbin Li, Xiangyu Yue, Chunlin Chen, Zhi Wang

reinforcement learningContinual reinforcement learningoffline reinforcement learninggenerative replaydiffusion models
72.60100
Fused
band ≈ ±15 pct pts (from σ = 0.30)
76.40100
Mimo
band ≈ ±22 pct pts (from σ = 0.44)
65.70100
DeepSeek
band ≈ ±20 pct pts (from σ = 0.40)

OpenReview ground truth

Rejected

TL;DR — We propose a dual generative replay framework that leverages expressive diffusion models to realize high-fidelity replay of the sample space for continual offline RL.

Abstract

We study continual offline reinforcement learning, a practical paradigm that facilitates forward transfer and mitigates catastrophic forgetting to tackle sequential offline tasks. We propose a dual generative replay framework that retains previous knowledge by concurrent replay of generated pseudo-data. First, we decouple the continual learning policy into a diffusion-based generative behavior model and a multi-head action evaluation model, allowing the policy to inherit distributional expressivity for encompassing a progressive range of diverse behaviors. Second, we train a task-conditioned diffusion model to mimic state distributions of past tasks. Generated states are paired with corresponding responses from the behavior generator to represent old tasks with high-fidelity replayed samples. Finally, by interleaving pseudo samples with real ones of the new task, we continually update the state and behavior generators to model progressively diverse behaviors, and regularize the multi-head critic in a behavior cloning manner to mitigate forgetting. Experiments on various benchmarks demonstrate that our method achieves better forward transfer with less forgetting, and closely approximates results of using previous ground-truth data due to its high-fidelity replay of the sample space.

Author context

Most prolific author: 6 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 34)