PapersWithELO
← ICLR 2024 leaderboard

Tackling Underestimation Bias in Successor Features by Distributional Reinforcement Learning

Mengxiao Lu, Yirui Zhou, Huojun Hong, Yaxin Peng, Xiaofeng Zhang, Yangchun Zhang

reinforcement learningSuccessor featuresDistributional reinforcement learningUnderestimation bias
28.30100
Fused
band ≈ ±15 pct pts (from σ = 0.30)
32.00100
Mimo
band ≈ ±22 pct pts (from σ = 0.43)
21.90100
DeepSeek
band ≈ ±21 pct pts (from σ = 0.43)

OpenReview ground truth

Rejected

Abstract

The framework of successor features (SFs) and generalized policy improvement (GPI) yields the potential to achieve zero-shot transfer in reinforcement learning (RL) among different tasks. However, GPI always suffers from inaccurate value function approximation in practice, resulting in a ``zero-shot'' somewhat fantastical. This paper focuses on comprehending the underlying causes of inaccurate SFs and presents a methodology for improving their accuracy. Our contributions encompass four key aspects: (i) we theoretically study the underestimation phenomenon in SF\&GPI; (ii) we introduce distributional RL into SF\&GPI, and demonstrate its effectiveness in relieving such underestimation; (iii) we show that distributional SFs (DSFs) is provided with a lower generalization bound than original SFs; (iv) we put forward that the performance of SFs-based algorithms can be enhanced by incorporating DSFs. Furthermore, we verify the quality of employing DSFs on the platform of multi-objective RL (MORL). Simulation study demonstrates the superiority of our concept in addressing underestimation challenges.

Author context

Most prolific author: 2 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 30)