Risk-Sensitive Variational Model-Based Policy Optimization
Alonso Granados, Jason Pacheco, Mohammadreza Ebrahimi
OpenReview ground truth
Abstract
RL-as-inference casts reinforcement learning (RL) as Bayesian inference in a probabilistic graphical model. While this framework allows efficient variational approximations it is known that model-based RL-as-inference learns optimistic dynamics and risk-seeking policies that can exhibit catastrophic behavior. By exploiting connections between the variational objective and a well-known risk-sensitive utility function we adaptively adjust policy risk based on the environment dynamics. Our method, $\beta$-VMBPO, extends the variational model-based policy optimization (VMBPO) algorithm to perform dual descent on risk parameter $\beta$. We provide a thorough theoretical analysis that fills gaps in the theory of model-based RL-as-inference by establishing a generalization of policy improvement, value iteration, and guarantees on policy determinism. Our experiments demonstrate that this risk-sensitive approach yields improvements in simple tabular and complex continuous tasks, such as the DeepMind Control Suite.
Author context
Most prolific author: 2 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 30 comparisons
Ranked above opponent in 52% of matchups.
- ▲ beat Learning Latent Structural Causal Models ×6
- ▲ beat Imagination Mechanism: Mesh Information Pr… ×4
- ▼ lost to Entropy-MCMC: Sampling from Flat Basins wi… ×4
- ▼ lost to A Unified Framework for Bayesian Optimizat… ×4
- ▼ lost to Forward $\chi^2$ Divergence Based Variatio… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 30)