Consistency Models as a Rich and Efficient Policy Class for Reinforcement Learning
Zihan Ding, Chi Jin
OpenReview ground truth
Abstract
Score-based generative models like the diffusion model have been testified to be effective in modeling multi-modal data from image generation to reinforcement learning (RL). However, the inference process of diffusion model can be slow, which hinders its usage in RL with iterative sampling. We propose to apply the consistency model as an efficient yet expressive policy representation, namely consistency policy, with an actor-critic style algorithm for three typical RL settings: offline, offline-to-online and online. For offline RL, we demonstrate the expressiveness of generative models as policies from multi-modal data. For offline-to-online RL, the consistency policy is shown to be more computational efficient than diffusion policy, with a comparable performance. For online RL, the consistency policy demonstrates significant speedup and even higher average performances than the diffusion policy.
Author context
Most prolific author: 3 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 38 comparisons
Ranked above opponent in 45% of matchups.
- ▲ beat Sparsity-Aware Grouped Reinforcement Learn… ×6
- ▲ beat Meta-Value Learning: a General Framework f… ×6
- ▲ beat Benchmarking Large Language Models as AI R… ×6
- ▼ lost to Achieving Minimax Optimal Sample Complexit… ×4
- ▼ lost to Expressive Modeling is Insufficient for Of… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 38)