PapersWithELO
← ICLR 2024 leaderboard

Meta-Value Learning: a General Framework for Learning with Learning Awareness

Tim Cooijmans, Milad Aghajohari, Aaron Courville

reinforcement learningmulti-agent reinforcement learningmeta-learning
20.00100
Fused
band ≈ ±15 pct pts (from σ = 0.30)
34.60100
Mimo
band ≈ ±22 pct pts (from σ = 0.44)
11.90100
DeepSeek
band ≈ ±21 pct pts (from σ = 0.42)

OpenReview ground truth

Rejected

TL;DR — We treat multi-agent optimization as a game and apply value-based RL to find nonexploitable cooperation in social dilemmas.

Abstract

Gradient-based learning in multi-agent systems is difficult because the gradient derives from a first-order model which does not account for the interaction between agents’ learning processes. LOLA (Foerster et al., 2018) accounts for this by differentiating through one step of optimization. We propose to judge joint policies by their long-term prospects as measured by the meta-value, a discounted sum over the returns of future optimization iterates. We apply a form of Q-learning to the meta-game of optimization, in a way that avoids the need to explicitly represent the continuous action space of policy updates. The resulting method, MeVa, is consistent and far-sighted, and does not require REINFORCE estimators. We analyze the behavior of our method on a toy game and compare to prior work on repeated matrix games.

Author context

Most prolific author: 5 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 32)