← ICLR 2024 leaderboard

Off-the-Grid MARL: Datasets with Baselines for Offline Multi-Agent Reinforcement Learning

Juan Claude Formanek, Asad Jeewa, Jonathan Phillip Shock, Arnu Pretorius

datasets & benchmarksreinforcement learningmulti-agent reinforcement learningmulti-agent systemsoffline reinforcement learningdatasets
19.20100
Fused
band ≈ ±14 pct pts (from σ = 0.29)
17.60100
Mimo
band ≈ ±20 pct pts (from σ = 0.41)
19.80100
DeepSeek
band ≈ ±20 pct pts (from σ = 0.40)

OpenReview ground truth

Rejected

TL;DR — Offline MARL is a nascent field which promises to turn large datasets into powerful, decentralised decision making systems. However, progress has been hampered by the lack of high-quality benchmark multi-agent datasets.

Abstract

Being able to harness the power of large datasets for developing cooperative multi-agent controllers promises to unlock enormous value for real-world applications. Many important industrial systems are multi-agent in nature and are difficult to model using bespoke simulators. However, in industry, distributed processes can often be recorded during operation, and large quantities of demonstrative data stored. Offline multi-agent reinforcement learning (MARL) provides a promising paradigm for building effective decentralised controllers from such datasets. However, offline MARL is still in its infancy and therefore lacks standardised benchmark datasets and baselines typically found in more mature subfields of reinforcement learning (RL). These deficiencies make it difficult for the community to sensibly measure progress. In this work, we aim to fill this gap by releasing off-the-grid MARL (OG-MARL): a growing repository of high-quality datasets with baselines for cooperative offline MARL research. Our datasets provide settings that are characteristic of real-world systems, including complex environment dynamics, heterogeneous agents, non-stationarity, many agents, partial observability, suboptimality, sparse rewards and demonstrated coordination. For each setting, we provide a range of different dataset types (e.g. Good, Medium, Poor, and Replay) and profile the composition of experiences for each dataset. We hope that OG-MARL will serve the community as a reliable source of datasets and help drive progress, while also providing an accessible entry point for researchers new to the field.

Author context

Most prolific author: 3 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Battle history — 34 comparisons

Ranked above opponent in 45% of matchups.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 34)