Off-the-Grid MARL: Datasets with Baselines for Offline Multi-Agent Reinforcement Learning
Juan Claude Formanek, Asad Jeewa, Jonathan Phillip Shock, Arnu Pretorius
OpenReview ground truth
TL;DR — Offline MARL is a nascent field which promises to turn large datasets into powerful, decentralised decision making systems. However, progress has been hampered by the lack of high-quality benchmark multi-agent datasets.
Abstract
Being able to harness the power of large datasets for developing cooperative multi-agent controllers promises to unlock enormous value for real-world applications. Many important industrial systems are multi-agent in nature and are difficult to model using bespoke simulators. However, in industry, distributed processes can often be recorded during operation, and large quantities of demonstrative data stored. Offline multi-agent reinforcement learning (MARL) provides a promising paradigm for building effective decentralised controllers from such datasets. However, offline MARL is still in its infancy and therefore lacks standardised benchmark datasets and baselines typically found in more mature subfields of reinforcement learning (RL). These deficiencies make it difficult for the community to sensibly measure progress. In this work, we aim to fill this gap by releasing off-the-grid MARL (OG-MARL): a growing repository of high-quality datasets with baselines for cooperative offline MARL research. Our datasets provide settings that are characteristic of real-world systems, including complex environment dynamics, heterogeneous agents, non-stationarity, many agents, partial observability, suboptimality, sparse rewards and demonstrated coordination. For each setting, we provide a range of different dataset types (e.g. Good, Medium, Poor, and Replay) and profile the composition of experiences for each dataset. We hope that OG-MARL will serve the community as a reliable source of datasets and help drive progress, while also providing an accessible entry point for researchers new to the field.
Author context
Most prolific author: 3 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 34 comparisons
Ranked above opponent in 45% of matchups.
- ▼ lost to WebArena: A Realistic Web Environment for … ×4
- ▼ lost to TiC-CLIP: Continual Training of CLIP Model… ×4
- ▼ lost to MMBench: Is Your Multi-modal Model an All-… ×4
- ▼ lost to Towards Assessing and Benchmarking Risk-Re… ×4
- ▲ beat RLP: A reinforcement learning benchmark fo… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 34)