2501.00160
Multi-Agent Reinforcement Learning couples agents' learning through a shared environment, producing emergent collective dynamics sensitive to initial conditions and parameters. Using a dynamical-syst…
A new discrete-time deterministic approximation model of independent multi-agent Boltzmann (softmax) Q-learning that explicitly accounts for agents' update frequencies, revealing that prior dynamical-systems approximations describe simplified variants rather than the original incremental algorithm. Applied to the Prisoner's Dilemma, it shows apparent long-term cooperation is a metastable, exploitable transient rather than a true equilibrium, and that raising the discount factor induces non-converging oscillations through a supercritical Neimark-Sacker bifurcation that converts the unique stable fixed point into an unstable focus surrounded by a stable limit cycle.
Multi-Agent Reinforcement Learning couples agents' learning through a shared environment, producing emergent collective dynamics sensitive to initial conditions and parameters. Using a dynamical-syst…