Logo image
LEARNING ZERO-SUM LINEAR QUADRATIC GAMES WITH IMPROVED SAMPLE COMPLEXITY AND LAST-ITERATE CONVERGENCE
Journal article   Peer reviewed

LEARNING ZERO-SUM LINEAR QUADRATIC GAMES WITH IMPROVED SAMPLE COMPLEXITY AND LAST-ITERATE CONVERGENCE

Jiduan Wu, Anas Barakat, Ilyas Fatkhullin and Niao He
SIAM journal on control and optimization, Vol.63(5), pp.3244-3271
01/01/2025

Abstract

Automation & Control Systems Mathematics Mathematics, Applied Physical Sciences Science & Technology Technology
Zero-sum linear quadratic (LQ) games are fundamental in optimal control and can be used (i) as a dynamic game formulation for risk-sensitive or robust control and (ii) as a benchmark setting for multiagent reinforcement learning with two competing agents in continuous state-control spaces. In contrast to the well-studied single-agent linear quadratic regulator problem, zero-sum LQ games entail solving a challenging nonconvex-nonconcave min-max problem with an objective function that lacks coercivity. Recently, Zhang et al. [Adv. Neural Inf. Process. Syst., 34 (2021), pp. 2949--2964] showed that an \epsilon-Nash equilibrium (NE) of finite-horizon zero-sum LQ games can be learned via nested model-free natural policy gradient algorithms with poly(1/\epsilon) sample complexity. In this work, we propose a simpler nested zeroth-order (ZO) algorithm improving sample complexity by several orders of magnitude and guaranteeing convergence of the last iterate. Our main result is twofold: (i) in the deterministic setting, we establish the first global last-iterate linear convergence result for the nested algorithm that seeks NE of zero-sum LQ games; (ii) in the model-free setting, we establish an \scrOilde\widet(\epsilon-2) sample complexity using a single-point ZO estimator. For our last-iterate convergence results, our analysis leverages the implicit regularization property and a new gradient domination condition for the primal function. Our key improvements in the sample complexity rely on a more sample-efficient nested algorithm design and a finer control of the ZO natural gradient estimation error utilizing the structure endowed by the finite-horizon setting.

Metrics

1 Record Views

Details

Logo image