Logo image
Heterogeneity, reinforcement learning, and chaos in population games
Journal article   Peer reviewed

Heterogeneity, reinforcement learning, and chaos in population games

Jakub Bielawski, Thiparat Chotibut, Fryderyk Falniowski, Michał Misiurewicz and Georgios Piliouras
Proceedings of the National Academy of Sciences - PNAS, Vol.122(25), p.e2319929121
24/06/2025
PMID: 40523171

Abstract

Applied Mathematics Economic Sciences
SignificanceThe emergence of chaotic behavior through coupled learning dynamics is a problem of fundamental importance that lies at the intersection of diverse fields such as social sciences, complexity science, mathematics, and artificial intelligence. Here, we provably show such a phenomenon in a setting where a large and diverse population of agents with strongly aligned interests learn concurrently. The collective dynamics can be provably chaotic, destabilizing the socially optimal equilibria and resulting in performance losses for all individuals and the society as a whole. Driving these results is a population-wide ergodic convergence where the time-average of the population-average behavior provably converges to its unique equilibrium value, despite the fact that the time-average behavior of any single agent may not converge. Inspired by the challenges at the intersection of Evolutionary Game Theory and Machine Learning, we investigate a class of discrete-time multiagent reinforcement learning (MARL) dynamics in population/nonatomic congestion games, where agents have diverse beliefs and learn at different rates. These congestion games, a well-studied class of potential games, are characterized by individual agents having negligible effects on system performance, strongly aligned incentives, and well-understood advantageous properties of Nash equilibria. Despite the presence of static Nash equilibria, we demonstrate that MARL dynamics with heterogeneous learning rates can deviate from these equilibria, exhibiting instability and even chaotic behavior and resulting in increased social costs. Remarkably, even within these chaotic regimes, we show that the time-averaged macroscopic behavior converges to exact Nash equilibria, thus linking the microscopic dynamic complexity with traditional equilibrium concepts. By employing dynamical systems techniques, we analyze the interaction between individual-level adaptation and population-level outcomes, paving the way for studying heterogeneous learning dynamics in discrete time across more complex game scenarios.
url
https://doi.org/10.1073/pnas.2319929121View
Published (Version of record) Open

Metrics

1 Record Views

Details

Logo image