Maximum Entropy On-Policy Actor-Critic via Entropy Advantage Estimation
Jean Seong Bjorn Choe, Jong-Kook Kim
OpenReview ground truth
TL;DR — We present a practical on-policy MaxEnt RL algorithm.
Abstract
Entropy regularisation is a widely adopted technique that enhances policy optimisation performance and stability. Many practical on-policy methods employ an entropy regularisation term to the policy gradient, thereby maximising policy entropy at visited states. On the other hand, another form of entropy regularisation, maximum entropy reinforcement learning (MaxEnt RL), augments the standard objective with an entropy term, aiming to maximise both the cumulative reward and the entropy of the trajectories induced by a policy. However, despite its empirical and theoretical achievements, its application in on-policy actor-critic contexts remains relatively underexplored. In this work, we propose an on-policy actor-critic algorithm based on the MaxEnt RL framework. A key aspect of our approach is separating the entropy objective from the MaxEnt RL objective. This delineation allows us to introduce an additional critic for the entropy objective alongside the conventional value critic. It also offers finer control over the optimisation process, incorporating a discount factor specifically for the entropy that provides a distinct way to balance the original and entropy objectives. Our empirical evaluations demonstrate that extending Proximal Policy Optimisation (PPO) and replacing its entropy regularisation with the proposed method significantly improves the performance of PPO in both continuous control tasks and across 16 Procgen environments. Additionally, the results underline MaxEnt RL's capacity to enhance generalisation.
Author context
Most prolific author: 1 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 34 comparisons
Ranked above opponent in 50% of matchups.
- ▼ lost to $\nu$-ensembles: Improving deep ensemble c… ×6
- ▲ beat Focus on Primary: Differential Diverse Dat… ×4
- ▲ beat HIPODE: Enhancing Offline Reinforcement Le… ×4
- ▼ lost to Language Model Inversion ×4
- ▲ beat On Trajectory Augmentations for Off-Policy… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 34)