ExplorerMathematicsMathematics
Research PaperResearchia:202608.10025

Wasserstein Policy Gradient for Entropy-Regularized Linear-Quadratic Control

Zhaoyu Zhu

Abstract

Wasserstein policy gradient (WPG) updates state-conditional action laws by transport in the action space. We study entropy-regularized discounted linear-quadratic (LQ) control. A Bellman verification argument shows that the unrestricted problem has a linear-Gaussian optimal policy, and the discounted-occupancy-weighted statewise Wasserstein gradient is tangent to this policy class. WPG therefore reduces exactly to a finite-dimensional ODE for the feedback gain and action covariance. We prove tha...

Submitted: August 10, 2026Subjects: Mathematics; Mathematics

Description / Details

Wasserstein policy gradient (WPG) updates state-conditional action laws by transport in the action space. We study entropy-regularized discounted linear-quadratic (LQ) control. A Bellman verification argument shows that the unrestricted problem has a linear-Gaussian optimal policy, and the discounted-occupancy-weighted statewise Wasserstein gradient is tangent to this policy class. WPG therefore reduces exactly to a finite-dimensional ODE for the feedback gain and action covariance. We prove that this ODE is globally well posed and converges exponentially from every admissible initialization. For each fixed LQ problem, the exponent has a positive limit as the entropy temperature tends to zero and contains no perturbative factor of the form exp(c/τ)\exp(-c/τ), while retaining the usual dependence on the conditioning of the control problem.


Source: arXiv:2608.07433v1 - http://arxiv.org/abs/2608.07433v1 PDF: https://arxiv.org/pdf/2608.07433v1 Original Link: http://arxiv.org/abs/2608.07433v1

Please sign in to join the discussion.

No comments yet. Be the first to share your thoughts!

Access Paper
View Source PDF
Submission Info
Date:
Aug 10, 2026
Topic:
Mathematics
Area:
Mathematics
Comments:
0
Bookmark