The Cyclical Chaos And Its Equilibrium
Stefan Juang, Yuan Zhou, Hong Wang, Elvis S. Liu, Nevin L. Zhang
OpenReview ground truth
Abstract
Finding a Nash Equilibrium (NE) in noncooperative games is a fundamental challenge in game theory and artificial intelligence, but existing methods can be computationally demanding to address the cyclical strategy problem. Existing methods like Policy Space Response Oracles (PSRO) allow agents to learn a best response (BR) policy against all prior policies. Once the learned policy converges, it is added to the sequence until an NE is identified. While the learning against all prior policies prevents agents' strategy interactions from descending into a cyclical chaos, this approach increases computational demands due to the expanding population of opponents. Our research offers a new perspective. We argue that cyclical strategies are not chaotic anomalies to be avoided; instead, they are orderly sequences integral to an equilibrium. We establish the theoretical equivalency between a complete set of cyclical strategies and the support set of a Mixed Strategy NE (MSNE). Our proof intuitively demonstrates that the cyclical strategies must form a circular counter, implying that a complete set is necessary to support an MSNE due to the intrinsic counterbalancing dynamic. This enables a novel graph search learning representation of self-play that finds an NE as a graph search. Our empirical results show improved self-play efficiency in discovering both a Pure Strategy NE (PSNE) and a MSNE in noncooperative games such as Connect4 and Naruto Mobile.
Author context
Most prolific author: 4 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 34 comparisons
Ranked above opponent in 44% of matchups.
- ▲ beat Learning Multiple Coordinated Agents under… ×6
- ▲ beat Sparsity-Aware Grouped Reinforcement Learn… ×6
- ▼ lost to Detecting Influence Structures in Multi-Ag… ×4
- ▲ beat From Images to Connections: Can DQN with G… ×4
- ▼ lost to Associative Transformer is a Sparse Repres… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 34)