Explorer›Mathematics›Mathematics
Research PaperResearchia:202610.05028

On the Convergence of Success Conditioning for Policy Optimization

Matthew Brun

Abstract

Success conditioning is a strategy for improving decision-making policies in stochastic environments; it updates a policy by increasing the probability of taking actions that yield successful outcomes. Success conditioning is common to many reinforcement learning applications, yet its limiting behavior and convergence rates are not well understood. In this work, we demonstrate that success conditioning converges to an optimal policy on a broad class of Markov decision processes (MDPs). We also d...

Submitted: October 5, 2026Subjects: Mathematics; Mathematics

Description / Details

Success conditioning is a strategy for improving decision-making policies in stochastic environments; it updates a policy by increasing the probability of taking actions that yield successful outcomes. Success conditioning is common to many reinforcement learning applications, yet its limiting behavior and convergence rates are not well understood. In this work, we demonstrate that success conditioning converges to an optimal policy on a broad class of Markov decision processes (MDPs). We also derive convergence rates in some common settings. For discounted MDPs, we prove convergence within O(1/εp)\mathcal{O}(1/\varepsilon^p) iterations to an ε\varepsilon-optimal policy, where the exponent pp depends on problem data. For single-period MDPs, such a policy is obtained within O(log⁡(1/ε))\mathcal{O}(\log(1/\varepsilon)) iterations.


Source: arXiv:2610.03642v1 - http://arxiv.org/abs/2610.03642v1 PDF: https://arxiv.org/pdf/2610.03642v1 Original Link: http://arxiv.org/abs/2610.03642v1

Please sign in to join the discussion.

No comments yet. Be the first to share your thoughts!

Access Paper
View Source PDF
Submission Info
Date:
Oct 5, 2026
Topic:
Mathematics
Area:
Mathematics
Comments:
0
Bookmark