RLP: A reinforcement learning benchmark for neural algorithmic reasoning
Yannick Niedermayr, Luca A Lanzendörfer, Benjamin Estermann, Roger Wattenhofer
OpenReview ground truth
TL;DR — Algorithmic reasoning is vital for problem-solving, while RL excels in certain tasks, its ability to handle complex algorithms is largely unexplored; we introduce an RL benchmark to assess this and find that RL struggles with algorithmic reasoning.
Abstract
Algorithmic reasoning is a fundamental cognitive ability that plays a pivotal role in problem-solving and decision-making processes. Although Reinforcement Learning (RL) has demonstrated remarkable proficiency in tasks such as motor control, handling perceptual input, and managing stochastic environments, its potential in learning generalizable and complex algorithms remains largely unexplored. To evaluate the current state of algorithmic reasoning in RL, we introduce an RL benchmark based on Simon Tatham's Portable Puzzle Collection. This benchmark contains 40 diverse logic puzzles of varying complexity levels, which serve as captivating challenges that test cognitive abilities, particularly in neural algorithmic reasoning. Our findings demonstrate that current RL approaches struggle with neural algorithmic reasoning, emphasizing the need for further research in this area. All of the software, including the environment, is available at https://github.com/rlppaper/rlp.
Author context
Most prolific author: 6 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 32 comparisons
Ranked above opponent in 41% of matchups.
- ▲ beat Imagination Mechanism: Mesh Information Pr… ×6
- ▲ beat Dispatching Ambulances using Deep Reinforc… ×6
- ▼ lost to WebArena: A Realistic Web Environment for … ×4
- ▼ lost to TiC-CLIP: Continual Training of CLIP Model… ×4
- ▼ lost to MMBench: Is Your Multi-modal Model an All-… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 32)