When May a Bandit Leave Its Anchor? E-Process-Authorized Thompson Sampling under Non-stationarity
Abstract
Stationarity rewards memory, but after a change the same history can mislead. We ask when forgetting should be permitted. E-process-authorized Thompson sampling (e-ATS) gives each arm full-history and discounted Beta states. An anytime-valid e-process first authorizes the discounted state, then a reversible relevance score controls its influence. Before authorization, e-ATS exactly follows optimistic Thompson sampling (OTS). Under a Beta-Bernoulli prior-predictive stationary model, e-ATS's proba...
Description / Details
Stationarity rewards memory, but after a change the same history can mislead. We ask when forgetting should be permitted. E-process-authorized Thompson sampling (e-ATS) gives each arm full-history and discounted Beta states. An anytime-valid e-process first authorizes the discounted state, then a reversible relevance score controls its influence. Before authorization, e-ATS exactly follows optimistic Thompson sampling (OTS). Under a Beta-Bernoulli prior-predictive stationary model, e-ATS's probability of ever departing from OTS is at most the chosen , without fitted thresholds. Relative to e-ATS, removing authorization increased mean normalized dynamic pseudo-regret by on the registered suite but reduced it by on the literature-derived replay suite. Therefore, evidence controls when adaptation begins, not whether it always helps.
Source: arXiv:2610.03646v1 - http://arxiv.org/abs/2610.03646v1 PDF: https://arxiv.org/pdf/2610.03646v1 Original Link: http://arxiv.org/abs/2610.03646v1
Please sign in to join the discussion.
No comments yet. Be the first to share your thoughts!
Oct 5, 2026
Data Science
Machine Learning
0