PapersWithELO
← ICLR 2024 leaderboard

Risk-Sensitive Variational Model-Based Policy Optimization

Alonso Granados, Jason Pacheco, Mohammadreza Ebrahimi

reinforcement learningReinforcement LearningVariational InferenceRisk Sensitive RLProbabilistic Inference
70.20100
Fused
band ≈ ±16 pct pts (from σ = 0.31)
75.30100
Mimo
band ≈ ±23 pct pts (from σ = 0.45)
66.40100
DeepSeek
band ≈ ±22 pct pts (from σ = 0.44)

OpenReview ground truth

Rejected

Abstract

RL-as-inference casts reinforcement learning (RL) as Bayesian inference in a probabilistic graphical model. While this framework allows efficient variational approximations it is known that model-based RL-as-inference learns optimistic dynamics and risk-seeking policies that can exhibit catastrophic behavior. By exploiting connections between the variational objective and a well-known risk-sensitive utility function we adaptively adjust policy risk based on the environment dynamics. Our method, $\beta$-VMBPO, extends the variational model-based policy optimization (VMBPO) algorithm to perform dual descent on risk parameter $\beta$. We provide a thorough theoretical analysis that fills gaps in the theory of model-based RL-as-inference by establishing a generalization of policy improvement, value iteration, and guarantees on policy determinism. Our experiments demonstrate that this risk-sensitive approach yields improvements in simple tabular and complex continuous tasks, such as the DeepMind Control Suite.

Author context

Most prolific author: 2 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Battle history — 30 comparisons

Ranked above opponent in 52% of matchups.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 30)