ExplorerData ScienceMachine Learning
Research PaperResearchia:202608.03068

Convergence and Regret of the Policy Gradient for Multi-Armed Bandits in Diffusion Environment

Yanwei Jia

Abstract

This paper studies the policy gradient update for a multi-arm bandit problem in diffusion environment that is described by a stochastic differential equation (SDE) under the continuous-time reinforcement learning framework by Wang et al. (2020), Jia and Zhou (2022b). With the logit parameterization for the stochastic policy, we show that it converges almost surely to the optimal arm under an arbitrary constant learning rate. Furthermore, we derive the non-asymptotic regret upper bound when the c...

Submitted: August 3, 2026Subjects: Machine Learning; Data Science

Description / Details

This paper studies the policy gradient update for a multi-arm bandit problem in diffusion environment that is described by a stochastic differential equation (SDE) under the continuous-time reinforcement learning framework by Wang et al. (2020), Jia and Zhou (2022b). With the logit parameterization for the stochastic policy, we show that it converges almost surely to the optimal arm under an arbitrary constant learning rate. Furthermore, we derive the non-asymptotic regret upper bound when the constant learning rate is below a time-invariant threshold; and the regret bound has order O(logT)O(\log T). We improve the analysis in Lattimore (2026a) for the same SDE by constructing a novel Lyapunov function and demonstrate the transparency of analyzing policy gradient using the tools in SDEs. In addition, the same Lyapunov function is also helpful in analyzing the discrete-time policy gradient algorithm.


Source: arXiv:2607.29593v1 - http://arxiv.org/abs/2607.29593v1 PDF: https://arxiv.org/pdf/2607.29593v1 Original Link: http://arxiv.org/abs/2607.29593v1

Please sign in to join the discussion.

No comments yet. Be the first to share your thoughts!

Access Paper
View Source PDF
Submission Info
Date:
Aug 3, 2026
Topic:
Data Science
Area:
Machine Learning
Comments:
0
Bookmark
Convergence and Regret of the Policy Gradient for Multi-Armed Bandits in Diffusion Environment | Researchia