Detecting Influence Structures in Multi-Agent Reinforcement Learning
Fabian Raoul Pieroth, katie.fitch@laralab.de, Lenz Belzner
OpenReview ground truth
TL;DR — Introducing novel, unified metrics and decentralized algorithms in MARL to precisely quantify agent influence, with empirical validation and convergence guarantees.
Abstract
We consider the problem of quantifying the amount of influence one agent can exert on another in the setting of multi-agent reinforcement learning (MARL). As a step towards a unified approach to express agents' interdependencies, we introduce the total and state influence measurement functions. Both of these are valid for all common MARL systems, such as the discounted reward setting. Additionally, we propose novel quantities, called the total impact measurement (TIM) and state impact measurement (SIM), that characterize one agent's influence on another by the maximum impact it can have on the other agents' expected returns and represent instances of impact measurement functions in the average reward setting. Furthermore, we provide approximation algorithms for TIM and SIM with simultaneously learning approximations of agents' expected returns, error bounds, stability analyses under changes of the policies, and convergence guarantees. The approximation algorithm relies only on observing other agents' actions and is, other than that, fully decentralized. Through empirical studies, we validate our approach's effectiveness in identifying intricate influence structures in complex interactions. Our work appears to be the first study of determining influence structures in the multi-agent average reward setting with convergence guarantees.
Author context
Most prolific author: 1 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 30 comparisons
Ranked above opponent in 60% of matchups.
- ▲ beat From Images to Connections: Can DQN with G… ×4
- ▲ beat Sparsity-Aware Grouped Reinforcement Learn… ×4
- ▼ lost to Meta-Value Learning: a General Framework f… ×4
- ▲ beat The Cyclical Chaos And Its Equilibrium ×4
- ▲ beat Learning Multiple Coordinated Agents under… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 30)