PapersWithELO
← ICLR 2024 leaderboard

Interpreting Categorical Distributional Reinforcement Learning: An Implicit Risk-Sensitive Regularization Effect

Ke Sun, Yingnan Zhao, Enze Shi, Yafei Wang, Xiaodong Yan, Bei Jiang, Linglong Kong

reinforcement learningdistributional reinforcement learningregularizationentropy
56.40100
Fused
band ≈ ±15 pct pts (from σ = 0.31)
53.50100
Mimo
band ≈ ±22 pct pts (from σ = 0.44)
62.00100
DeepSeek
band ≈ ±21 pct pts (from σ = 0.43)

OpenReview ground truth

Rejected

TL;DR — We interpret distributional reinforcement learning from the perspective of regularization effect.

Abstract

The theoretical advantages of distributional reinforcement learning~(RL) over expectation-based RL remain elusive, despite its remarkable empirical performance. Starting from Categorical Distributional RL~(CDRL), our work attributes the potential superiority of distributional RL to its \textit{risk-sensitive entropy regularization}. This regularization stems from the additional return distribution information regardless of only its expectation via the return density function decomposition, a variant of the gross error model in robust statistics. Compared with maximum RL that explicitly optimizes the policy to encourage the exploration, we reveal that the resulting risk-sensitive entropy regularization of CDRL plays a different role as an augmented reward function. It implicitly optimizes policies for a risk-sensitive exploration towards true target return distributions, which helps to reduce the intrinsic uncertainty of the environment. Finally, extensive experiments verify the importance of this risk-sensitive regularization in distributional RL, as well as the mutual impacts of both explicit and implicit entropy regularization.

Author context

Most prolific author: 5 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 32)