Meta-Value Learning: a General Framework for Learning with Learning Awareness
Tim Cooijmans, Milad Aghajohari, Aaron Courville
OpenReview ground truth
TL;DR — We treat multi-agent optimization as a game and apply value-based RL to find nonexploitable cooperation in social dilemmas.
Abstract
Gradient-based learning in multi-agent systems is difficult because the gradient derives from a first-order model which does not account for the interaction between agents’ learning processes. LOLA (Foerster et al., 2018) accounts for this by differentiating through one step of optimization. We propose to judge joint policies by their long-term prospects as measured by the meta-value, a discounted sum over the returns of future optimization iterates. We apply a form of Q-learning to the meta-game of optimization, in a way that avoids the need to explicitly represent the continuous action space of policy updates. The resulting method, MeVa, is consistent and far-sighted, and does not require REINFORCE estimators. We analyze the behavior of our method on a toy game and compare to prior work on repeated matrix games.
Author context
Most prolific author: 5 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 32 comparisons
Ranked above opponent in 39% of matchups.
- ▼ lost to Expressive Modeling is Insufficient for Of… ×6
- ▲ beat Consistency Models as a Rich and Efficient… ×6
- ▼ lost to Achieving Minimax Optimal Sample Complexit… ×4
- ▲ beat Detecting Influence Structures in Multi-Ag… ×4
- ▼ lost to Pre-training with Synthetic Data Helps Off… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 32)