DiffCPS: Diffusion Model based Constrained Policy Search for Offline Reinforcement Learning
Longxiang He, Linrui Zhang, Junbo Tan, Xueqian Wang
OpenReview ground truth
TL;DR — We propose a novel approach called Diffusion Model based Constrained Policy Search (DiffCPS), which tackles the diffusion-based constrained policy search without resorting to Advantage Weighted Regression.
Abstract
Constrained policy search (CPS) is a fundamental problem in offline reinforcement learning, which is generally solved by advantage weighted regression (AWR). However, previous methods may still encounter out-of-distribution actions due to the limited expressivity of Gaussian-based policies. On the other hand, directly applying the state-of-the-art models with distribution expression capabilities (i.e., diffusion models) in the AWR framework is insufficient since AWR requires exact policy probability densities, which is intractable in diffusion models. In this paper, we propose a novel approach called $\textbf{Diffusion Model based Constrained Policy Search (DiffCPS)}$, which tackles the diffusion-based constrained policy search without resorting to AWR. The theoretical analysis reveals our key insights by leveraging the action distribution of the diffusion model to eliminate the policy distribution constraint in the CPS and then utilizing the Evidence Lower Bound (ELBO) of diffusion-based policy to approximate the KL constraint. Consequently, DiffCPS admits the high expressivity of diffusion models while circumventing the cumbersome density calculation brought by AWR. Extensive experimental results based on the D4RL benchmark demonstrate the efficacy of our approach. We empirically show that DiffCPS achieves better or at least competitive performance compared to traditional AWR-based baselines as well as recent diffusion-based offline RL methods. Code will be made publicly available upon acceptance.
Author context
Most prolific author: 7 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 28 comparisons
Ranked above opponent in 55% of matchups.
- ▲ beat Look Ma, No Training! Observation Space De… ×4
- ▲ beat Exploring the State and Action Space in Re… ×4
- ▼ lost to The Update-Equivalence Framework for Decis… ×4
- ▲ beat CLIP as Multi-Task Multi-Kernel Learning ×4
- ▲ beat Optimal criterion for feature learning of … ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 28)