Tackling Underestimation Bias in Successor Features by Distributional Reinforcement Learning
Mengxiao Lu, Yirui Zhou, Huojun Hong, Yaxin Peng, Xiaofeng Zhang, Yangchun Zhang
OpenReview ground truth
Abstract
The framework of successor features (SFs) and generalized policy improvement (GPI) yields the potential to achieve zero-shot transfer in reinforcement learning (RL) among different tasks. However, GPI always suffers from inaccurate value function approximation in practice, resulting in a ``zero-shot'' somewhat fantastical. This paper focuses on comprehending the underlying causes of inaccurate SFs and presents a methodology for improving their accuracy. Our contributions encompass four key aspects: (i) we theoretically study the underestimation phenomenon in SF\&GPI; (ii) we introduce distributional RL into SF\&GPI, and demonstrate its effectiveness in relieving such underestimation; (iii) we show that distributional SFs (DSFs) is provided with a lower generalization bound than original SFs; (iv) we put forward that the performance of SFs-based algorithms can be enhanced by incorporating DSFs. Furthermore, we verify the quality of employing DSFs on the platform of multi-objective RL (MORL). Simulation study demonstrates the superiority of our concept in addressing underestimation challenges.
Author context
Most prolific author: 2 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 30 comparisons
Ranked above opponent in 41% of matchups.
- ▼ lost to Optimal Sample Complexity for Average Rewa… ×4
- ▼ lost to Efficient Action Robust Reinforcement Lear… ×4
- ▼ lost to Revisiting the Static Model in Robust Rein… ×4
- ▲ beat Interpreting Categorical Distributional Re… ×4
- ▲ beat CLIP-Guided Reinforcement Learning for Ope… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 30)