Interpreting Categorical Distributional Reinforcement Learning: An Implicit Risk-Sensitive Regularization Effect
Ke Sun, Yingnan Zhao, Enze Shi, Yafei Wang, Xiaodong Yan, Bei Jiang, Linglong Kong
OpenReview ground truth
TL;DR — We interpret distributional reinforcement learning from the perspective of regularization effect.
Abstract
The theoretical advantages of distributional reinforcement learning~(RL) over expectation-based RL remain elusive, despite its remarkable empirical performance. Starting from Categorical Distributional RL~(CDRL), our work attributes the potential superiority of distributional RL to its \textit{risk-sensitive entropy regularization}. This regularization stems from the additional return distribution information regardless of only its expectation via the return density function decomposition, a variant of the gross error model in robust statistics. Compared with maximum RL that explicitly optimizes the policy to encourage the exploration, we reveal that the resulting risk-sensitive entropy regularization of CDRL plays a different role as an augmented reward function. It implicitly optimizes policies for a risk-sensitive exploration towards true target return distributions, which helps to reduce the intrinsic uncertainty of the environment. Finally, extensive experiments verify the importance of this risk-sensitive regularization in distributional RL, as well as the mutual impacts of both explicit and implicit entropy regularization.
Author context
Most prolific author: 5 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 32 comparisons
Ranked above opponent in 53% of matchups.
- ▼ lost to Optimal Sample Complexity for Average Rewa… ×4
- ▲ beat CLIP-Guided Reinforcement Learning for Ope… ×4
- ▼ lost to Efficient Action Robust Reinforcement Lear… ×4
- ▲ beat Focus on Primary: Differential Diverse Dat… ×4
- ▲ beat HIPODE: Enhancing Offline Reinforcement Le… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 32)